01
Edition
05
picks
The machines that check the machine's work showed up.
The machines that check the machine's work showed up. Friday we flagged an unshipped edition — Forall's proofs, Sentinel's code-aware QA, Libretto's self-healing tests, all launching into the exhaustion Pydantic's essay named — and said the theme completes when the cluster re-trends or a major vendor ships a verification gate. One business day later the vendors shipped past the startups. Vercel Labs' deepsec trended: coding agents auditing your whole repo, with a PR-diff mode built for CI gating. Replay.io launched Replay QA: an agent that explores your app, records what broke, and files root-caused reports for your coding agent to fix. Under both sits the evidence infrastructure — crabbox, from the OpenClaw org, records auditable receipts for every remote test run, and LoRA Speedrun refuses to list a fine-tuning record until it's been re-run three times on frozen hardware. The demand side supplied its own arithmetic: exploit brokers pay $500,000 for WordPress RCEs, and a researcher reports finding one with GPT5.6 and $25 in credits — dropped as an exploit writeup, kept as the reason your scanner now needs to be an agent too. And the day's loudest story, the report that Claude Fable produced a counterexample to the Jacobian Conjecture (433 points), is news for exactly one reason: a counterexample doesn't ask you to trust the model, only to check its output. That's the whole theme. Also dropped per rubric: Inkling's 975B open-weights model launch, and Orion, Kagi's browser.
02
Replay QA — finds what broke before your users do
03
crabbox — warm a box, sync the diff, run the suite
04
LoRA Speedrun — records don't count until re-run
05
One of these,
every weekday.
Free. Unsubscribe by replying with one word. No tracking pixels in the email.