# astra comes in at fable's price, and the four endpoints fell together

> openai ships astra at the same api price as fable 5.1, cost per completed task becomes the metric of the day, and in the afternoon the major providers went down almost together

- edition: Thursday, September 3, 2026 (2026-09-03)
- notebook: dev & ai
- topics: llm · market · papers · brazil
- items: 11 from 13 sources
- original: https://tonho.wtf/en/daily/2026-09-03/
- portuguese edition: https://tonho.wtf/diario/2026-09-03/
- authorship: written by an llm pipeline, reviewed and translated by antonio leandro (tonho.wtf)

---

The whole day was quoting the same product: an hour of machine work. OpenAI launched Astra at Fable 5.1's exact list price — $10 per million input, $50 per million output — which takes the fight off the "who is smarter" axis and puts it on "what does the finished task cost". Latent Space converted that into dollars per hour, the Pragmatic Engineer reports companies cutting half the bill by moving simple load to open models, and Nvidia paid $12.93bn for the place where those open models live. Four independent items, one question: what is the unit price of the work, and who owns the infrastructure that delivers it.

The afternoon answered the second half of the question in an uncomfortable way. OpenAI, Claude, Gemini and Grok showed up down at almost the same time, with no public cause so far, and the most interesting hypothesis to come out of the thread is that interchangeability is the mechanism: when one falls, everybody migrates at the same moment and takes the others down. Worth rereading in light of Tuesday, when the thread was that the token got cheaper and the task got 20% more expensive: today Astra claims exactly the opposite on the agentic axis, with cost per task below half of Fable 5's for the same score, according to Artificial Analysis. It's a third-party claim, not a self-report — and even so it's a claim, not the experience of someone who put it in production.

## labs

**[GPT‑6 Astra](https://simonwillison.net/2026/Sep/3/gpt6-astra/)** — gradual rollout to ChatGPT Plus, Pro, Business and Enterprise, plus API and AWS, with the label `gpt-6-astra`. Simon Willison hasn't tested it yet and says so; what he brings is a reading of the numbers. The 99.9% on ARC‑AGI 3 comes with a big asterisk: it was obtained for $19k with OpenAI's own "Provider Adapter harness", which preserves opaque reasoning state between requests and does compaction to reuse earlier work, while the benchmark's standard harness gave 62.7% for $26k. That is almost 40 points of difference that aren't in the model, they're in the harness — the same lesson that showed up here at the end of August. On security the jump is big (100% on ExploitBench against GPT‑5.6 Sol's 78.5%, 42.4% on ExploitGym against 30.3%), and on long context the internal eight-needle benchmark gives 100% between 256K and 512K tokens and 96.3% between 512K and 1M. On Artificial Analysis's Intelligence Index, though, it ties with Sol at 61 and sits five points below Fable 5.1.

**[Claude Commerce Agents](https://gihyo.jp/article/2026/09/claude-commerce-agents?utm_source=feed)** — Anthropic has released under Apache 2.0 a blueprint with reference implementations of two agents, one for buying and one for running a store, with demos in retail, travel, telecom and ticket sales, plus a Claude Code plugin called `commerce-builder`. The architecture detail is what matters: a main agent holds the context (cart, history, preference) instead of handing each domain off to a subagent, and the safeguards don't live only in the prompt — there are runtime gates that check target, limit and approval state at tool-call time, with a price or stock change held until a human approves. The gains cited (cart up to 35% bigger, 60% more chance of completing the purchase) come with no company and no measurement conditions. The code is a reference and won't get ongoing maintenance.

## research

**[NeoMME](https://arxiv.org/abs/2609.01657)** — family of bidirectional encoders at 260M and 800M parameters that processes multilingual text and raw image patches in a single transformer, trained from scratch with a masked discrete diffusion objective, 16,384-token context (enough for two 4K UHD images). The thesis is economic: visual document retrievers like ColPali reuse a generative VLM and carry the parameter and compute overhead of an architecture built to generate text on a task that generates nothing. On ViDoRe v3, the 260M retriever gets 0.523 nDCG@10 and the 800M reaches 0.556; on an L40S with 2048x2048 input, the 260M encodes pages at about 2x the throughput of ColModernVBERT. Hierarchical token pooling plus asymmetric quantization compress the late interaction embeddings by 255x while preserving more than 95% of base nDCG@10. Checkpoints under Apache 2.0 and a contribution to Transformers.

**[The Vocabulary Gap Is an Equity Gap](https://arxiv.org/abs/2609.01645)** — the most useful counter-example of the day for anyone evaluating RAG. The benchmark has 51 federal benefit eligibility rules and 25 information needs, each written in two registers (the agency's, and that of someone asking in plain or non-native English), with the correct passage fixed. On BM25, TF‑IDF and a term-graph retriever, the formal register gives Recall@5 of 96% to 100%; the informal one collapses to 36% to 44%. On BM25, Recall@1 drops from 84% to 16%. The mechanism is measured, not inferred: the formal query shares 0.63 of its content terms with the gold passage, the informal one only 0.11. A simple, auditable plain-to-formal lexicon takes Recall@5 from 44% to 80%. If your retrieval evaluation uses questions written by the people who wrote the documents, it is measuring the easy case — and that holds just the same in Portuguese, where public-agency jargon and the user's actual question never touch.

## brazil

**[Vayou](https://www.tabnews.com.br/ohgawa/vayou-player-de-video-nativo-que-traduz-a-legenda-sozinho-rust-slint)** — video player for Windows and Linux in Rust + Slint, presented on TabNews, rewritten from scratch after a first version in Tauri + Svelte that spent browser-sized RAM to draw controls on top of the video. The render trick is the part worth reading: mpv runs with `vo=libmpv` and draws into the same OpenGL framebuffer Slint uses, so it's one window, one swapchain and no frame taking a walk through the CPU. The trap he reports is good too: on Linux, if the render context doesn't get the X11/Wayland display, libmpv doesn't open a VA display and falls back to software decoding silently, no error, with the CPU pinned as the only symptom. Subtitle translation pulls the track with ffmpeg, translates in blocks preserving ASS formatting and hands it back as an external track — via an unofficial Google Translate endpoint, and the app warns you when a stretch wasn't translated instead of faking it. MIT, 8.8 MB binary, fat Windows installer because it ships libmpv and ffmpeg with it.

## market

**[Nvidia buys Hugging Face for $12.93bn](https://www.heise.de/news/Nvidia-uebernimmt-Hugging-Face-fuer-12-93-Milliarden-US-Dollar-11440209.html?wt_mc=rss.red.ho.ho.atom.beitrag.beitrag)** — binding agreement signed on Wednesday, with about $11.9bn for shareholders and up to $1bn in a share program for those who move over, closing expected in the first half of 2027 subject to regulatory approval. The promise that the platform stays open and keeps supporting other chip makers is in the blog post and also in the SEC filing, which changes how much it weighs. In the risk section of that same filing, Nvidia registers what actually threatens the business: a good share of the most popular open models comes from China, and regulatory restriction on open models would limit the hub's supply. Worth remembering who is buying: Nvidia has already published more than 500 models and 250 datasets there and invested in the company in 2023.

**[Ask HN: why did OpenAI, Claude and Grok go down at the same time](https://news.ycombinator.com/item?id=49551096)** — 327 points and no public cause. The thread takes apart part of its own premise: Down Detector is self-reported and uses report volume against a baseline, so people rushing to check whether a service is down already produce the spike — and there's a report of someone using Gemini with no failure at all in the same window. What's left that's most concrete is the displaced-demand hypothesis: since the products are perceived as interchangeable, one going down pushes users to the others within minutes. If that's the mechanism, the cross-provider fallback strategy a lot of people wrote this year is exactly the cause of the problem it was supposed to solve.

## world

**[I refused to train the AI that could replace me](https://restofworld.org/2026/ai-training-jobs-expert-replacement/?utm_source=rss&utm_medium=rss&utm_campaign=feeds)** — James Maisiri writes in Rest of World about the invitation, fresh out of a PhD, to train a system to design assessments, teach undergraduates and grade essays: 600 rand ($37) an hour, in a country where the minimum wage is 30.23 rand ($2) an hour and youth unemployment was 47.4% in the second quarter of 2026. The point of the piece isn't the labor-market number, it's what is being bought: not the knowledge, but the judgment — how you decide an essay is worth 75% and not 60%. Platforms like Outlier, Mercor and Surge are recruiting doctors, lawyers and engineers to transfer exactly that kind of discretion. The detail that closes the argument: the 45-minute interview was conducted by an AI, which afterwards emailed the candidate their strengths and weaknesses and suggested redoing part of the assessment.

## who wrote

**[an automated AI Engineer you can hire for <$6 an hour](https://www.latent.space/p/astra)** — Latent Space got early access and burned more than 20 billion Astra tokens on real tasks, reporting managing a fleet of subagents, reading logs, instrumenting and debugging an entire system in one go. The number in the title is explicit arithmetic: 33 tokens per second at the top rate of $50 per million. That's throughput, not completed work — and the piece itself warns that in Ultra mode, running in parallel, spend blows well past it (one of the tasks described cost about $100 over two days). The text is incomplete by their choice, with a promise to finish it later.

**[The Pulse: tech companies move to open AI models](https://newsletter.pragmaticengineer.com/p/the-pulse-tech-companies-move-to)** — only the teaser is available (paid edition). What it announces: Uber, Pinterest, Stripe, Coinbase, Ramp and AT&T cutting about 50% of the AI bill by swapping proprietary models for open models on simple load, with model routing. If the number holds up in the body of the text, it's the data point that ties the day to the Hugging Face acquisition.

**[The call for an agentic standard](https://uxdesign.cc/the-call-for-an-agentic-standard-we-need-to-stop-shipping-the-same-form-four-different-times-be5b9d40e370)** — also truncated, only the teaser: we need to stop shipping the same form four times. The thesis announced is that, with the agent becoming the client of the interface, repetition stops being a question of aesthetics and becomes a question of standards.

## stalled sources

Anthropic Engineering at 102 days without publishing, Karpathy at 126, Interconnects at 17, Sebastian Raschka at 13, Anthropic Research at 7.
