01
AI Hacker Daily
Today
06
picks
# AI Hacker Daily — 2026-08-28 Every tool today exists because a written claim is not evidence.
02
RealDiff — runs the tests on both sides of the PR and diffs what the functions returned
03
Restoredrill — restores the backup into a throwaway container and writes the auditor's JSON
04
Opslane — runs the fix in a sandbox before the PR exists, and its AGENTS.md lists the ways its own gate lies
05
experiential — a 0%-markup gateway that turns your traffic into a router by simulating the runs it never made
06
One of these,
every weekday.
Free. Unsubscribe by replying with one word. No tracking pixels in the email.
Archive
2026-08-27
Today's tools are sorted by how much of the company is still inside them
6 picks
2026-08-26
The approval click is being re-engineered, not removed
4 picks
2026-08-25
The stack is losing parts, and the state is moving into a bucket. Three projects on the same front page today run one architecture — object storage is the only source of truth and every server is a disposable cache: a git host that went from zero to 1,341 stars in under two days and carries Shopify's CEO as its author, a Kafka-shaped stream server with "zero-disk architecture" on its landing page, and an ebook library that is a Cloudflare Worker over an R2 bucket with no database. The why-now was named in the walgit thread by the one commenter who noticed the coincidence: S3 shipped compare-and-swap 21 months ago, "and it has just been absolutely wild the massive flocking towards disaggregated storage" since. We watched the same primitive reach the Durable Objects layer on August 6 (celld: state in your own S3 bucket); today it reaches git, streams, and a bookshelf. The other three picks delete a different part by the same instinct — a container runtime that is one 1.52 MB static binary with no daemon and zero RAM at rest, a Claude Code memory tool whose README argues a shared memory service is a single point of failure and puts the index in a per-project SQLite file, and a personal search engine that is one binary over your own browsing history, from the person who built searx. The six are ordered by what got deleted: the database, the broker's disks, the app server, the daemon, the memory service, the cloud. The counterweight is in the same thread as the headline pick: the sharpest walgit comment is not about architecture, it is a distributed-systems engineer reading the S3 lease code and finding a HEAD-compare-DELETE race, in a repository whose README was accused two comments later of being written by a model. A bucket makes the state durable. It does not make the code correct.
6 picks
2026-08-24
The coding harness just became a component you swap, not a home you live in. An essay called "What Is a Harness?" hit the front page this weekend (27 points) proposing the equation Agent = Model + Harness — the model is rented, the harness is the part you can own, adapt, and point at a different model when the pricing changes. The top skeptic in the thread called "harness" "the AI hype word for 2026 after agent in 2025," and he may be right about the word while the thing ships around him, because today's pool is what commoditization looks like when it arrives as artifacts instead of arguments: a skill that lets Claude Code dispatch Codex or Grok as a disposable second opinion; a config compiler that declares your agent setup once, lockfile-pins it, and materializes it per harness; an IDE that runs Claude Code, Codex, OpenCode, Cursor and Grok side by side in isolated worktrees; a menu-bar app for the person who lost track of which worktree the agents are in; and a Rust MCP server that gives any harness eyes and hands on a real Windows desktop. Around the edges, the aftermarket that forms when a part standardizes: three separate skill-and-plugin directories trending on the same day — one at 31,597 stars indexing "1000+ agent skills" across five harnesses, one an official Anthropic community-marketplace mirror, one a 3,332-star hub. Three days ago the best sentence in our biggest thread was a developer explaining why he runs every model's output past every other model: "The tokens are too cheap not to." Today that sentence is an installable skill. The five picks are ordered by how far the harness has drifted from being singular: consulted, configured, parallelized, watched, embodied.
5 picks
2026-08-21
The expensive component is not the one doing the work: a 1990s keyword baseline outscores four hyperscalers on EnterpriseRAG-Bench, a Bedrock config gap bills $1,182 for a cache never read, and today's picks are ordered by how much of the expensive part you can delete.
5 picks
2026-08-20
The approval click lost, so the agent is being handed its own computer: six projects ordered by how far the box sits from the machine in front of you
6 picks
2026-08-19
AUTH is a leash, not a wall: we reproduced the guardrail benchmark (74% detected, 8% blocked) and ordered the picks by who is left holding the decision
6 picks
2026-08-18
Benchmarks got cheap to generate: grading picks by whether the number could have come out badly
6 picks
2026-08-17
The layer in the middle, and whether it tells you what it takes
5 picks
2026-08-14
The agent loop went free this week; the evidence did not — five stack layers open-sourced, ordered by who published what they measured
5 picks
2026-08-13
The token accounting layer: every rung measures the bill more precisely and none of them stops it
5 picks