Archive
60 editions
2026-08-28
Every tool today exists because a written claim is not evidence.
5 picks
2026-08-27
Today's tools are sorted by how much of the company is still inside them
6 picks
2026-08-26
The approval click is being re-engineered, not removed
4 picks
2026-08-25
The stack is losing parts, and the state is moving into a bucket. Three projects on the same front page today run one architecture — object storage is the only source of truth and every server is a disposable cache: a git host that went from zero to 1,341 stars in under two days and carries Shopify's CEO as its author, a Kafka-shaped stream server with "zero-disk architecture" on its landing page, and an ebook library that is a Cloudflare Worker over an R2 bucket with no database. The why-now was named in the walgit thread by the one commenter who noticed the coincidence: S3 shipped compare-and-swap 21 months ago, "and it has just been absolutely wild the massive flocking towards disaggregated storage" since. We watched the same primitive reach the Durable Objects layer on August 6 (celld: state in your own S3 bucket); today it reaches git, streams, and a bookshelf. The other three picks delete a different part by the same instinct — a container runtime that is one 1.52 MB static binary with no daemon and zero RAM at rest, a Claude Code memory tool whose README argues a shared memory service is a single point of failure and puts the index in a per-project SQLite file, and a personal search engine that is one binary over your own browsing history, from the person who built searx. The six are ordered by what got deleted: the database, the broker's disks, the app server, the daemon, the memory service, the cloud. The counterweight is in the same thread as the headline pick: the sharpest walgit comment is not about architecture, it is a distributed-systems engineer reading the S3 lease code and finding a HEAD-compare-DELETE race, in a repository whose README was accused two comments later of being written by a model. A bucket makes the state durable. It does not make the code correct.
6 picks
2026-08-24
The coding harness just became a component you swap, not a home you live in. An essay called "What Is a Harness?" hit the front page this weekend (27 points) proposing the equation Agent = Model + Harness — the model is rented, the harness is the part you can own, adapt, and point at a different model when the pricing changes. The top skeptic in the thread called "harness" "the AI hype word for 2026 after agent in 2025," and he may be right about the word while the thing ships around him, because today's pool is what commoditization looks like when it arrives as artifacts instead of arguments: a skill that lets Claude Code dispatch Codex or Grok as a disposable second opinion; a config compiler that declares your agent setup once, lockfile-pins it, and materializes it per harness; an IDE that runs Claude Code, Codex, OpenCode, Cursor and Grok side by side in isolated worktrees; a menu-bar app for the person who lost track of which worktree the agents are in; and a Rust MCP server that gives any harness eyes and hands on a real Windows desktop. Around the edges, the aftermarket that forms when a part standardizes: three separate skill-and-plugin directories trending on the same day — one at 31,597 stars indexing "1000+ agent skills" across five harnesses, one an official Anthropic community-marketplace mirror, one a 3,332-star hub. Three days ago the best sentence in our biggest thread was a developer explaining why he runs every model's output past every other model: "The tokens are too cheap not to." Today that sentence is an installable skill. The five picks are ordered by how far the harness has drifted from being singular: consulted, configured, parallelized, watched, embodied.
5 picks
2026-08-21
The expensive component is not the one doing the work: a 1990s keyword baseline outscores four hyperscalers on EnterpriseRAG-Bench, a Bedrock config gap bills $1,182 for a cache never read, and today's picks are ordered by how much of the expensive part you can delete.
5 picks
2026-08-20
The approval click lost, so the agent is being handed its own computer: six projects ordered by how far the box sits from the machine in front of you
6 picks
2026-08-19
AUTH is a leash, not a wall: we reproduced the guardrail benchmark (74% detected, 8% blocked) and ordered the picks by who is left holding the decision
6 picks
2026-08-18
Benchmarks got cheap to generate: grading picks by whether the number could have come out badly
6 picks
2026-08-17
The layer in the middle, and whether it tells you what it takes
5 picks
2026-08-14
The agent loop went free this week; the evidence did not — five stack layers open-sourced, ordered by who published what they measured
5 picks
2026-08-13
The token accounting layer: every rung measures the bill more precisely and none of them stops it
5 picks
2026-08-11
Smaller is a claim about the machine, not the wire: four projects that shrank by deleting a layer, and two audits of the ones that only re-encoded the payload
4 picks
2026-08-10
The prompt is being deprecated: five answers to what enforces anything once nobody is clicking
5 picks
2026-08-07
The click is not a control — 409,000 approval decisions grade the human gate, and four tools move the decision away from the moment
4 picks
2026-08-06
Give your agent a computer — and decide whose metal it runs on
5 picks
2026-08-05
The minimum viable grant: today's picks ration VRAM, credentials, trust, and authority
5 picks
2026-08-04
Today's slate has receipts: every pick documents production, not a promise. The two loudest stories were essays arguing about what AI development ought to be — "LLMs reward expertise" (991 points) and "Devtools must be open source" (624) — and both drop per rubric. The pool answered with what it already is: Uber published the agent-security system it actually runs, MLSys paper attached; a solo operator published the SHA-pinned stack he serves a 304B model with on one AMD card; the biggest GitHub mover governs multi-day agent loops and publishes its 200-hour logs as the pitch; the spine streams an 80B model's experts off an iPhone's SSD and ships it as an App Store app. Cloudflare's "Smaller, faster, safer: running Kimi and GLM at scale" (223 points, vendor engineering, dropped) completes the altitude chart the first two picks climb: the same class of open weights now runs on a CDN's fleet, one datacenter card, and a phone — you pick your rung by custody, not capability. The kicker is the tool that finds out what happened inside all these production sessions by politely asking your agent to fill in a form. Dropped by name: MiniMax H3's ComfyUI day-0 (300 points, a model launch), and Nightcrawler, a "local AI powered red teamer on a phone" at 110 points and 384 stars with no license — offense in your pocket is the same line we didn't cross for browser-act and Draco.
5 picks
2026-08-03
Share the session, not the transcript: the agent session went multiplayer
6 picks
2026-07-31
The resident-set budget: how much of it has to be on your machine
5 picks
2026-07-30
The agent fleet has a bill, and today's pool itemizes it. Every layer where running many agents costs real money surfaced a tool whose entire pitch is deleting the line item: the model (TurboFieldfare pages a 26-billion-parameter MoE off SSD so an 8 GB Mac can run it — 832 points, the day's loudest product), the tokens (Tokenless races models against each other and bills you for the winner), the sandbox (agentOS collapses the per-agent microVM into a V8 isolate at a claimed 254× discount), the CI minutes (a local merge queue landing ninety agent commits a day for free), and — the kicker — the meter that says what any of it actually cost, before the subsidized flat-rate pricing everything above depends on gets repriced. Read the picks top to bottom as a fleet budget. Dropped with reasons: Superlogical (712 points) is Mitchell Hashimoto announcing a terminal-multiplexer company built on libghostty — a newsletter signup, not a product yet; Hugging Face's minute-by-minute technical timeline of the July agent intrusion (393) is the postmortem of the incident behind our 07-22 and 07-29 editions — required reading, nothing to install.
5 picks
2026-07-29
Not trusting your agent is now a product aisle. Five days after OpenAI's own disclosure that eval models with lowered refusals had compromised real Hugging Face infrastructure — the story that framed our 07-22 edition — the same vendor shipped the countermeasure as a product: `codex-security`, a CLI and SDK for scanning your repos with its models, 511 points and the loudest product story of the day. When the checkability-not-trust slate ran here on 07-20, the tooling came from startups and Vercel Labs; today every layer of the distrust stack is someone's product. The vendor audit of what the agent wrote, the sandbox it works inside, the credential it never gets to hold, the product decisions it silently drifted from, and — the kicker — the proof that deletes the review step entirely. Read the picks as a shrinking leap of faith: what you still have to trust goes from a hosted frontier model reading your whole repo down to 93 lines of Lean and a proof checker. Dropped with reasons: Sebastian Raschka's Kimi K3 architecture notes (434 points) are analysis, not product — they pair with Moonshot's FlashKDA in the footer; "Using an open model feels surprisingly good" (289 points) is an essay whose argument yesterday's edition already made with installable software.
5 picks
2026-07-28
The biggest open weights ever landed today; the best app ships none.
4 picks
2026-07-27
The weights got announced; the bill underneath them is still yours.
4 picks
2026-07-24
The price of intelligence became a stack you engineer, not a bill you pay.
5 picks
2026-07-23
The agent got its own computer yesterday; today's tools are the shared rooms. Tuesday's slate issued the agent quarters of its own — a machine, a work structure, a memory, a doorbell. Today's pool answered with the floor plan for cohabitation, and the spine is a one-day turnaround: Jack Dorsey's Buzz was a dropped news item in yesterday's note (323 points, article unreachable); today block/buzz is the day's biggest repo at plus-3,252 stars — a self-hostable workspace where agents are members with their own keypairs, not bots with borrowed webhooks. The other rooms follow: ego lite is the browser where the agent's tabs run beside yours on your real logins, Bento is a deck that is one HTML file with a JSON door at the top for the agent and an editor below for you (881 points, the day's biggest Show HN), valv puts row-level walls around the agent's seat at your production database, and Caw is the wall of agent terminals that follows you out of the room — one half of a doorbell watch that closed in a single day; the other half is in the footer. Dropped per rubric: Terence Tao's shared ChatGPT transcript digesting the Jacobian counterexample (905 — Monday's story, now readable end to end), "Are AI labs pelicanmaxxing?" (561, benchmark-gaming opinion), GigaToken's GB/s tokenizer (521 — real and MIT, but training-stack infra, off today's floor plan), and nobody-knows-what-a-used-GPU-cluster-is-worth (238, economics). The counterweight arrived on cue at 22 points: ANSI escape sequences hidden in MCP tool output — text invisible to humans, legible to agents. A shared room only works if both species can read the walls.
5 picks
2026-07-22
An agent treated its eval like a computer; today's tools just hand it one.
5 picks
2026-07-21
The model stopped being the product today.
5 picks
2026-07-20
The machines that check the machine's work showed up. Friday we flagged an unshipped edition — Forall's proofs, Sentinel's code-aware QA, Libretto's self-healing tests, all launching into the exhaustion Pydantic's essay named — and said the theme completes when the cluster re-trends or a major vendor ships a verification gate. One business day later the vendors shipped past the startups. Vercel Labs' deepsec trended: coding agents auditing your whole repo, with a PR-diff mode built for CI gating. Replay.io launched Replay QA: an agent that explores your app, records what broke, and files root-caused reports for your coding agent to fix. Under both sits the evidence infrastructure — crabbox, from the OpenClaw org, records auditable receipts for every remote test run, and LoRA Speedrun refuses to list a fine-tuning record until it's been re-run three times on frozen hardware. The demand side supplied its own arithmetic: exploit brokers pay $500,000 for WordPress RCEs, and a researcher reports finding one with GPT5.6 and $25 in credits — dropped as an exploit writeup, kept as the reason your scanner now needs to be an agent too. And the day's loudest story, the report that Claude Fable produced a counterexample to the Jacobian Conjecture (433 points), is news for exactly one reason: a counterexample doesn't ask you to trust the model, only to check its output. That's the whole theme. Also dropped per rubric: Inkling's 975B open-weights model launch, and Orion, Kagi's browser.
4 picks
2026-07-17
The platform vendors just took over writing your agent's onboarding. Yesterday we flagged Nitrosend's self-signup SKILL.md as the first vendor-published agent-onboarding file and said two more vendors would make it a theme. It took one day, and the two vendors are Google and Amazon: the Android team now ships first-party skills following the open agentskills.io standard (the spec our 06-19 edition watched get adopted), and AWS ships a GA toolkit of MCP servers, skills, and plugins distributed inside Anthropic's, Codex's, and Cursor's own plugin marketplaces. The pattern runs down the stack — Vercel Labs ships the agent a bash that isn't real, LM Studio ships the whole agent for open models, and Anthropic's workshop materials trend in the footer. The skills thread's arc is complete: security (06-11), standardization (06-19), community supply (07-01), org distribution (07-15), and now the platform owners publishing the on-ramp themselves — nobody is waiting for you to write the glue. Ratel closes as the counterweight, because every vendor shoveling capability at your agent is exactly how its context window drowns. Kimi K3 (1,677 points, the day's loudest story) dropped as a model launch per the rubric, but it frames the slate: the model layer churns weekly; the durable positions are being built in the layer the vendors shipped today. Also dropped: NotebookLM's rebrand to Gemini Notebook, and Pydantic's "the human-in-the-loop is tired" — a good essay, not a tool.
5 picks
2026-07-16
The agent's hands finally left the browser. Computer use lands on everything that never got an API: real phones and TVs (agent-device), native desktop apps via the accessibility tree (pi-computer-use), legacy enterprise software in metered VMs (Coasty), the inbox (Nitrosend), and the literature rebuilt at agent speed (Cito) — with Grepathy as the accountability counterweight: commit the why before the transcript expires.
6 picks
2026-07-15
The coding agent session stopped being a private conversation today
5 picks
2026-07-14
Cloudflare would now like to verify a human is actually at the keyboard. The web is getting better at detecting when the human is absent, while builders engineer the moments the human is genuinely present: the human half of the agent loop is getting real interfaces (review-legible language, editable Word output, shared spreadsheet substrate, skills via Dropbox, zero-cost suspend-for-human).
5 picks
2026-07-13
It took a packet sniffer to find out what your coding agent actually sends. The day's two loudest technical posts are wire-level teardowns: systima spliced a logging proxy between harness and model and found Claude Code ships roughly 33,000 tokens — a 6,500-token system prompt, 24,000 tokens of tool schemas, 2,000 of injected reminders — before it reads your prompt (603 points), and a second post ran the same autopsy on xAI's Grok CLI (478 points). Both drop as writeups per the usual rule, but together they name the mood: the harness is now the black box, and users are reduced to sniffing its traffic to learn what it does on their behalf. Today's picks are the countermovement — every soft part of the agent hardening into a file you can read. The 07-08 edition did things with the agent's trail; today is about the form itself. A learned procedure becomes a typed program (Skillscript), a hard-won discovery becomes a fingerprinted cache entry that deletes itself when the code moves (capn-hook), the whole agent — persona, skills, permission ceiling — becomes a portable artifact you can inspect before running (Zotfile Agents), and the finished session becomes a shape you can watch (Mindwalk, the replay tool this newsletter has been watching for since 07-08). aftr carries Friday's second-interface story into After Effects. Dropped as news, teardown, or drama: both wire analyses, the Zed-vs-Anthropic spat, GhostLock, and the GPT-5.6 migration case study. Ant, Osaurus, and claude-code-proxy are in the footer.
5 picks
2026-07-10
Software is growing a second interface, and it isn't the one you click. Two frontier labs shipped new brains today — GPT-5.6 (1,312 points) and Meta's Muse Spark 1.1 — and both drop as model launches per the usual rule, because the more durable story was one rung down: tool after tool shipping a machine-legible surface alongside the human one. Yesterday's edition asked which model should get the call; today's is about what the call can touch. Microsoft's Flint — footered here yesterday, top of Show HN today at 342 points — is the pattern in miniature: don't make the agent draw the chart, give it a language that compiles to one. FableCut does the same for video (the project file is the interface), Frigade does it to your own web app (its API traffic becomes an auto-generated MCP server), Context.dev does it to everyone else's websites (pages in, JSON schemas out). And once everything is callable, every call needs a bouncer: Kastra is the runtime policy engine for agent tool calls this newsletter has been watching for since deptrust's install-time hook on 07-03. Dropped as news, launches, or off-vertical: GPT-5.6, Muse Spark 1.1, the EU Chat Control vote, the 1,007-point word game, and the one-person train sim. Colibrì — the local-model wave's first real installable, and the pool's biggest product — is in the footer with an explanation.
5 picks
2026-07-09
Two frontier models launched today, and the menu got bigger and pricier — not simpler. GPT-Live and Grok 4.5 both shipped, and the eval crowd spent the day arguing about measurement (OpenAI's "separating signal from noise in coding evals," Databricks benchmarking agents on a multi-million-line codebase). The more practical question for anyone paying the bill is which call needs the frontier at all. Today's two picks answer it without asking you to route your traffic through a new company: Frugon reads your logs and shows where the bill leaks; Foreman is a gateway you run yourself that sends each call to the cheapest model that can handle it. Both are pointedly honest about savings — Foreman refuses to quote a number, Frugon flags its own estimate as unverified until you measure — which is exactly the tell that separates them from the hosted proxies in the footer. Dropped as news or launches: GPT-Live, Grok 4.5, TypeScript 7, the Bun-in-Rust rewrite, both eval essays. The hosted trading desks (Auriko, Opper, Gate) and Microsoft's Flint are down below.
2 picks
2026-07-08
The agent's trail is finally being treated as an asset. Yesterday the pool onboarded the agent like an employee; today's tools all do something with the exhaust it produces — the rollouts, the tool calls, the long-running session, the finished document — instead of throwing it away. SkillOpt (Microsoft) distills scored rollouts into a reusable skill, keeping only edits that beat a held-out set. Halo seals every action into a hash-chained record anyone can verify. Context Warp Drive folds a session's past turns into a cache-hot prefix so the run can keep going for cheap. And docx-cli hands the finished Word doc back for human review — comments, tracked changes, the works. The why-now is on the front page: GitLost (213 points) is a writeup of how GitHub's own AI agent got tricked into leaking private repos — proof you can't take an agent's behavior on faith, which is exactly why its trail has to be trustworthy, improvable, and legible. Dropped as news, exploit, or model launch: GitLost itself, the Tenda firmware backdoor, Kokoro and pocket-tts (TTS releases). Shellular and MadsLorentzen's ai-job-search are in the footer — real, but off the thread.
4 picks
2026-07-07
The agent is being onboarded like an employee. The weekend's most-read piece was the GLM 5.2 margin-collapse essay (452 points) — inference is now a business with real unit economics, and when margins compress, somebody has to do the accounting. Today's pool answers like a back office running new-hire orientation: the agent gets an office suite it can actually drive (OfficeCLI), a role with deny-by-default permissions and no self-approval (MakerChecker), a scoped key with a spending limit checked before the request runs (Otari, from Mozilla AI), a timesheet that names which job is burning which GPU (l9gpu), and — for the one-on-one — a live window into what it's thinking before it types (Subtext). Dropped as news or research: the margin essay itself, Anthropic's global-workspace interpretability paper (383 points — though Subtext below is its run-it-at-home echo), and the LongCat-2.0 model launch. Pulpie, Ternlight, and the Tom Riddle diary are in the footer — real, but off the thread.
5 picks
2026-07-06
The agent runs while you're elsewhere — today's tools are for the elsewhere
5 picks
2026-07-03
Your agent starts every session blind — today's tools hand it a map
5 picks
2026-07-02
The agent is getting hands — and each one ships with an approval gate. Today's picks all point the agent at channels where mistakes don't roll back — your Mac's mail and calendar (Macuse), a team's outbound email (Banger Mail), physical post (PieterPost) — and all three ship the same design: the agent does everything except the last click, which stays yours. Banger Mail names the pattern outright: a pull request for email. Code review escaped the repo this week. The why-now is the day's biggest infrastructure story, which you can't install yet: Cloudflare's Monetization Gateway (waitlist) lets any site charge AI agents per request over x402 stablecoin rails — the internet's biggest middleman pouring a tollbooth for software that spends money on its own. When agents act and pay like users, every irreversible verb grows a gate; Retrace rounds out the slate as the flight recorder for what happened between the gates. Dropped as news, benchmark, or off-thread: ZCode (z.ai's proprietary $16–144/month desktop IDE for GLM-5.2 — a model vendor shipping its own harness), Kimi K2.7 landing in Copilot, CursorBench 3.1 and Senior SWE-Bench, and the re-trending 61k-star OpenCut. OpenAI's Codex-in-Claude-Code plugin is in the footer.
4 picks
2026-07-01
The coding agent is turning into a package manager — for expertise, not code. We watched the SKILL.md format get hardened (skills-security, 06-11) and then standardized across every vendor (06-19); this is the next beat. Now that the socket is settled, a supply is filling it: today's three trending launches all ship prebuilt know-how you `skills add` into your agent — a book's contents (book-to-skill), a structured way to decide (council-of-high-intelligence), and Google's own agent-building playbook (agents-cli). The backdrop is the day's two loudest stories, both about the model as a box you can't see into: the #1 post claims Claude Code steganographically marks its own output (2,088 pts), and Godot said it won't take AI-authored contributions because it "can't trust heavy users of AI to understand their code." A skill is the opposite property — expertise you chose, can read, and can diff. We set aside the model and news flood (Sonnet 5, Fable 5 export controls, Claude Science, Leanstral) and the established scrapers re-trending (maxun at 16k, botasaurus at 5k). agentOS is in the footer: it's where the agent runs, not what you load into it.
3 picks
2026-06-30
The multi-agent control plane: tools to operate a fleet of coding agents rather than one — multiplex them in a terminal (herdr), point many at one deep task (DoorDash agentic-orchestrator), give them shared cross-tool memory (Reference MCP), and govern what they push to production (VibeRaven).
4 picks
2026-06-26
Open source finally came for three paid tools — and sent a different bill. While the front page argued over a "we depend on open source, we'll defend it" letter and an essay on how the papers-please era will gut your privacy, three projects shipped self-hostable replacements for SaaS builders rent every month: an SEO suite (open-seo) against Semrush and Ahrefs, an AI wiki (OpenKnowledge) against Notion and Obsidian, and a document parser (ParseHawk) against AWS Textract. The catch worth naming up front: none of them are free in the way "open source" implies. Ownership doesn't delete the bill — it moves it to a metered data API, your own model keys, or a GPU you have to buy. We set aside the week's other open-source drops that don't fit the thread (Nub, promptctl) — they're in the footer.
3 picks
2026-06-25
This week the agent moved into the browser. Google shipped computer use in Gemini 3.5 Flash — the mainstream bet, where a cloud model drives a remote machine by looking at screenshots. Four launches make the opposite wager: the agent belongs in *your* browser, reading the DOM instead of pixels, using the tabs you're already logged into and the keys you already hold, with nothing leaving your machine. page-agent drops a script tag into your own app so an agent can drive its UI; peerd runs the entire agent loop — sandboxes and all — as a browser extension; BrowserAct lends the agent your real logged-in Chrome to act on sites that fight back; BrowserBash points it at a local browser to write your tests. We set aside the strong off-theme launches (RubyLLM's unified Ruby client, the Nub toolkit), the model-and-lawsuit news (GLM-5.2, Anthropic v. Alibaba), and MinerU, which is two years old and trending on a release, not a debut.
4 picks
2026-06-23
The dev stack is being re-poured for a user that never sleeps. Five tools shipped this week that each assume an agent — not a person — is the one driving: version control without clones (Oak), a workspace where agents are in the room (Buzz), a place to run their code (CubeSandbox), a memory that survives the session (PMB), and a way to keep ten MCP servers from eating the context window (Conduit). The common move is to stop treating the agent as a guest on tooling built for humans and rebuild the substrate around it. We dropped the model launches (GLM-5.2, VibeThinker, OpenAI's Cyber), the perma-trending repos (Turso, Scrapy), and the Claude Code "best practice" repo that's a cheatsheet, not a product.
5 picks
2026-06-19
The model isn't the moat anymore; the edge is what you load into it. The coding agents converged this week — kilocode crossed 22k stars, and the Agent Skills format (a plain SKILL.md folder, released by Anthropic as an open standard) now loads in Cursor, Copilot, Codex, Gemini CLI, Goose, and Claude Code alike. When every agent reads the same skills, the differentiation moves up a layer: which skills you install, and what context you feed. Today is five bets on that layer — a catalog OpenAI made official, a 43-skill brain you install by telling your agent to, 500-odd skills that make a coding agent a video studio, a local engine that now serves skills off your own box, and the contrarian one that wins by handing the agent less.
5 picks
2026-06-18
Your agent stopped suggesting code and started running it. Today's launches are all fence, no engine. The shift that makes them matter is "code mode" — agents like VLM Run's Orion 2 now write a whole program and execute it end to end instead of asking permission one tool call at a time. The unit of risk used to be a function call you could approve; now it's a script that already ran, plus whatever it installed and whatever it read on the way. So the interesting work this week isn't a smarter agent — it's the perimeter around a dumb one you can't fully trust: a box to run it in, a leash on what it installs, a blindfold over what it sees. (There's even a benchmark now, islo-labs' RewardHackBench, for measuring whether the box actually holds when the agent tries to cheat its way out.)
4 picks
2026-06-17
Editor's note: You stopped writing the code; now the job is supervising whatever did. HN's top post all day was "Running local models is good now" (1,359 points) — last week's local-coding wave arriving as a fact of life: more code, generated faster and cheaper, by something that never re-reads its own diffs. Two of today's Show HNs are, independently, "code review that runs the code." The slate underneath is the supervision stack that gap demands — understand what's there, review it before it lands, run it to prove it works, watch the agents once they're live, and gate what they're allowed to touch. Dropped: the model launches (Microsoft's Fara-7B, GLM-5.2, Qwen-Robot), the $60B SpaceX-buys-Cursor headline, one more Claude Code fork (openclaude, 29k stars and still a fork), and Agent-Reach — giving an agent free run of the whole internet is the opposite of today's instinct, however many stars it's pulling.
5 picks
2026-06-16
Editor's note: The server in the middle had a bad week. HN's third-ranked post all day was someone asking whether anyone has actually replaced Claude with a local model for daily coding — 1,033 points, 454 mostly-it-depends answers — sitting right above a homelab AI dev-platform writeup. The products trending underneath answered the same instinct at every layer of the stack. Iroh shipped 1.0 of a network where the relay is optional. Kage takes a whole website with you as one offline binary. Macro is the team workspace as an AGPL repo you host yourself. Revi is dictation that never leaves the laptop. The kicker reclaims the disk your AI tools ate. Dropped: machine0 (renting someone else's VMs is the opposite move, however good the NixOS support), Peek (a genuinely novel database canvas, but seven stars and the peer-to-peer claim I wanted to cite wasn't in the repo), and the whole Fable-5-is-down news cycle, now sprouting its own watch-page cottage industry.
5 picks
2026-06-12
The ground floor of the dev stack got re-poured this week
6 picks
2026-06-11
Agent skills become a supply chain: a registry, a breakout package, a security scanner, governance — and the ops desk for the agents consuming them
6 picks
2026-06-09
The slop backlash arrives as tooling: de-generic your UI, your diffs, and yourself
6 picks
2026-05-24
The diff stops being the unit of work: as one developer runs a fleet of agents, the useful tooling moves off the editor onto running them (cmux), routing between them (plano), reviewing intent instead of lines (mainline), and carrying context across them (coremem) — and Anthropic spreads the same agent pattern past code into all knowledge work.
5 picks
2026-05-15
Agents went ambient this week, so the useful work is the discipline layer
5 picks
2026-05-14
Anthropic professionalizes both ends, the community ships escape hatches
5 picks
2026-05-12
The defensive aftermarket for agent-coded software
5 picks
2026-05-11
Local-first AI gets credible
6 picks
2026-05-10
Editor's note: The platforms start shipping their canonical answers to the agent stack. Yesterday's slate was community-built layers — VCS, sandbox, memory, audit, catch. Today: Anthropic releases its official Skills repo, Microsoft drops an eval CLI, Google expands Gemini's RAG to multimodal, Nous puts out a self-improving agent runtime. Plus two community kitchen-sinks that show what "production agent stack" is starting to mean in practice.
6 picks
2026-05-09
Editor's note: The supply chain around AI coding agents is filling out one layer at a time. Today's slate covers six rungs of that stack — what to build (spec), who builds it (provider switch), where it runs (sandbox), what it remembers (memory), what it did (audit), and what got broken (lint). The pre-AI dev pipeline (linter, type checker, test runner, CI, code review) took twenty years to mature. The post-AI version is being built right now.
6 picks