01
Edition
07
picks
# AI Hacker Daily — 2026-08-19 We ran the benchmark we kept saying we would run, and it returns 74 against 8.
# AI Hacker Daily — 2026-08-19 We ran the benchmark we kept saying we would run, and it returns 74 against 8. Doberman is the agent-security library this newsletter called the best evidence document it had ever graded, one day ago, on the strength of its documentation alone. So this morning we cloned it cold, installed it, and ran the suite: 137 labeled rows, 112 attacks, no model called, no API key, about four minutes. Every cell matches what they published — overall TPR **0.74**, and `tpr_strict` **0.08**. Read that pair slowly. Three quarters of the attacks are caught; one in twelve is actually stopped. Everything in between is an `AUTH` — a prompt, handed to a person, to decide. Their own docs have the phrase for it: "AUTH is a leash, not a wall." And the leash is attached to somebody who is measurably busier than they were two years ago. Linear published its aggregate product data the same morning: pull requests up **111%** across 47,900 paid workspaces since June 2024, teams running coding agents going from 21 to 65 PRs a week while teams without them went 8 to 10, and roughly **48%** of new issues now authored by AI. Linear is admirably blunt about what that is not — the report calls it "motion rather than value," notes that time spent on existing tasks simply held while AI became "a new layer of work," and states plainly: "We have no way of knowing whether this increased output led to positive business outcomes." So: three times the output, the same day, and the safety story for most of it ends in a dialog box. Today's picks are ordered by one question — when the check runs, who is left holding the decision? The list starts with tools where the answer is nobody, and ends with one where the answer is a model that publishes no accuracy number at all.
02
keychain-store — the security check that never asks you anything
03
Edgemetry — no consent banner, because there is nothing to consent to
04
Shoehorn — quantization where a solver decides, and prints what the fit cost
05
Claude Watermark — the detector that refuses to give you the verdict you want
06
Argus — the bottom rung: the check is a model, and there is no number on it
07
One of these,
every weekday.
Free. Unsubscribe by replying with one word. No tracking pixels in the email.