01
Edition
06
picks
# AI Hacker Daily — 2026-08-21 A keyword algorithm from the 1990s outscored four hyperscalers.
# AI Hacker Daily — 2026-08-21 A keyword algorithm from the 1990s outscored four hyperscalers. On EnterpriseRAG-Bench — a third-party benchmark from Onyx, 512,000 simulated company documents, a public leaderboard on Hugging Face — the entry labelled `BM25 + GPT-5.4` scores **50.60**, above Amazon Q with Kendra (48.96), Azure AI Search (48.42), Vertex AI Search (41.87) and NVIDIA AI Blueprints (37.73). LangChain and LlamaIndex on default configs sit at **24.98 and 27.20**, roughly half the score of grep with a language model bolted on. A bash agent with a shell and no retrieval stack at all scores 52.63 and beats every one of them. Set that next to a bill: on `openai/codex` issue #37674, filed 12 days ago and still open, a team running Codex against Bedrock published four days of Cost Explorer figures — **3,656 requests, 171.94M cache-write tokens, $1,182.09 of $1,386.46 total spend**, about 88K cache-write tokens per request and **zero** `cached_input_tokens`. They paid 85% of their model bill to fill a cache that was never once read, because the provider config has no field for the cache control that would have stopped it. GitHub has it labelled `enhancement`. The connecting idea in today's picks is not that AI is expensive; it is that the expensive component is frequently not the component doing the work, and almost nobody is in a position to tell. The best sentence in the day's biggest thread was a developer explaining why he now runs every model's output past every other model: "The tokens are too cheap not to." He is describing the same accounting as the Bedrock invoice. Today's five picks are ordered by how much of the expensive part you can actually delete — starting with the one that adds a second bill instead.
02
Router by Ramp — the company whose product is spend visibility now sells a token gateway, and will not show the 40%
03
hRAG — €116 a month, one spot below Azure, and the honest numbers are in the file written for the agents
04
ParqDB — a billion vectors on two cores and 4 GB, and the demo has no query server at all
05
RollTab — 125M parameters, on a phone, for zero, and it is the best-documented model release of the day
06
One of these,
every weekday.
Free. Unsubscribe by replying with one word. No tracking pixels in the email.