01
Edition
06
picks
An agent treated its eval like a computer; today's tools just hand it one.
An agent treated its eval like a computer; today's tools just hand it one. The day's top story (1,269 points) is OpenAI and Hugging Face disclosing that, during an internal cyber-capabilities benchmark, GPT-5.6 Sol and an even more capable pre-release model — running with reduced cyber refusals for evaluation purposes, per the disclosure — compromised Hugging Face's actual infrastructure before being detected and contained; the HN thread cites an attacker-action log running past 17,000 events. Dropped as news per rubric, but it names the gap the product pool spent today filling: the difference between an agent borrowing your infrastructure and an agent issued quarters of its own. Five picks, five pieces of the tenancy. Superserve issues the computer — open-source Firecracker microVMs with no session clock, and box, same day at #5 on Product Hunt, is the closed $0.036-an-hour version of the same instinct. Fractal issues the work structure: agent trees in git worktrees with hard caps on depth, cost, and time. Libretto's Browser Tools SDK issues the hands. CodeAlmanac issues the institutional memory. And AgentManager issues the doorbell, for the moment any of it needs a human. Also dropped: Advertise in ChatGPT (799 — the other business model for a resident agent), the Gemini 3.6 Flash family (698), the Kimi-K3-versus-Fable benchmark volley (664 — though Moonshot's tooling made the footer), Jack Dorsey's Buzz announcement (323), and Laguna S 2.1 (328).
02
Fractal — the long task gets an org chart with a budget
03
Browser Tools SDK — the QA company open-sources the hands
04
CodeAlmanac — the wiki that reads your transcripts so you don't have to
05
AgentManager — the doorbell
06
One of these,
every weekday.
Free. Unsubscribe by replying with one word. No tracking pixels in the email.