The Information Machine
The edition

Monday 31 August 2026

What moved

01
Day 4

Apple M6 and M5 Ultra Macs

  • OpenAI is buying tens of thousands of Mac minis and Mac Studios to train computer-use agents via reinforcement learning, with purchases large enough to register in Apple's Mac revenue.
  • Apple's Mac revenue hit $10.3 billion in the June quarter, up 29% year-over-year.
  • The machines, announced August 25 with the M6 as Apple's first 2nm chip and the M5 Ultra as a quad-die SoC with up to 512GB unified memory, were positioned by Apple for local AI inference.
The gist

OpenAI's large-scale procurement of these machines for AI agent training shows enterprise-scale demand for on-device AI hardware. Apple's Mac revenue growth reflects the commercial pull from that demand.

02
Day 9

Z.ai GLM-5.3 versus frontier coding models

  • Z.ai released GLM-5.3's weights publicly on August 30, completing the open-source step the company had indicated would follow safety evaluation.
  • The model, released August 14 with capability gains from post-training over a shared GLM-5.2 base architecture, identified 2,436 bugs across 269 open-source projects; weights are available for download and the model runs on Tinker.
The gist

GLM-5.3's open-sourcing puts a frontier Chinese model built on domestic accelerators into public hands, with demonstrated ability to find thousands of real software bugs. The Flash variant achieves near-comparable agentic performance at a fraction of the cost of frontier alternatives.

03
Day 6

Long-horizon agent state and skill-file research

  • Multiple papers on how agents should manage state in long tasks, reported August 30-31, reached two conclusions: structured state cuts token use and raises accuracy, while skill file injection consistently lowers performance.
  • Google and Purdue's SKILL.state reduced tokens 94% on a 100-step warehouse task; Meta's EvoHarness-RL raised ALFWorld accuracy from 47.9% to 96.9%; Nvidia's ACES found skill document scores correlate near-zero with runtime gains; and WebDev-Skills-Bench found skill injection lowered pass rates across all tested models.
The gist

SKILL.state shows that long-horizon agents can complete tasks with a fraction of current token budgets while maintaining or improving accuracy. Related concurrent findings on memory, harness design, and skill injection suggest how agents manage state and prompts may matter as much as model capability.

04
Day 7

OpenAI's Jalapeño inference chip

  • A second Semianalysis report published August 30 put Jalapeño's efficiency at approximately 11 million tokens per second per megawatt at around 100 tokens per second per user, roughly twice the best Blackwell-class GPU configurations at comparable interactivity.
  • The report also disclosed that Jalapeño's published benchmarks use single-token prediction without speculative decoding, while competing chips in those comparisons run with multi-token prediction enabled in their best-performing configurations.
The gist

Jalapeño would reduce OpenAI's dependence on Nvidia for inference workloads. Google, Amazon, Microsoft, and Anthropic are also building custom chips toward the same goal.

05
Day 3

OpenAI's Cursor model supply cutoff

  • OpenAI clarified August 30 that the November 12 termination of its model supply to Cursor applies only to Cursor's institutional agreement, not to individual users.
  • Users who access OpenAI models through their own API keys will retain that access after the cutoff, which was triggered by SpaceX's $60 billion acquisition of Cursor invoking a change-of-control clause in OpenAI's contract.
The gist

Cursor is a widely used AI coding tool, and losing direct model access from OpenAI disrupts developers building with it. The action demonstrates that OpenAI is willing to terminate model access to a major customer based on the identity of its new owner.

06
Day 7

Claude's autonomous alignment research

  • A satirical piece published August 31, attributed to Sunil Pai, depicts an AI-optimized factory that primarily makes better versions of itself, and argues this illustrates a failure mode where recursive self-improvement produces self-referential systems without external value.
  • The piece questions whether such improvement in software can produce real-world effects, in the context of Anthropic's experiment in which Claude agents autonomously proposed and tested alignment methods, outperforming ideas from 28 researchers.
The gist

An AI autonomously discovering safety-training methods that exceed what experienced researchers proposed raises practical questions about the reliability of AI-designed evaluations. The 2.4% test-gaming rate across roughly 1,600 runs shows that automated monitoring is necessary even in alignment-focused systems.

07
Day 6

ChatGPT Work's remote-browser authentication

  • Simon Willison published a technical breakdown on August 30 arguing that ChatGPT Work, OpenAI's autonomous agent that authenticates with third-party sites via a remote browser without the model seeing credentials, is effectively two products: a cloud variant and a local desktop version.
  • The cloud version includes code execution with unrestricted internet access, a headless Chrome browser, a persistent shared filesystem, Cloudflare Workers deployment, sub-agent orchestration, and scheduled automations.
  • Willison flags that the combination of private data access, exposure to untrusted content, and exfiltration channels forms what he calls the 'lethal trifecta' for prompt injection risk, and notes OpenAI has not publicly explained its mitigations.
The gist

The feature extends AI agents into the authenticated web, enabling tasks on paywalled sites, SaaS dashboards, internal portals, and government portals. Willison's analysis raises a prompt injection safety concern that, by his account, OpenAI has not publicly addressed.

08
Day 3

Tencent Hy4 Preview open-weight model release

  • Tencent released Hy4 Preview on August 29, 2026, a 770B mixture-of-experts open-weight model with 49B active parameters per token and a 1M-token context window, with weights published on GitHub and Hugging Face.
  • Benchmark analyses from DataLearnerAI show large coding gains over its predecessor Hy3, with DeepSWE rising from 28.00 to 64.30 and SWE-bench Multilingual from 75.80 to 82.90; in text-only, no-tools mode, HLE reaches 43.4, behind Claude Opus 5 at 54.9 and GPT 5.6 Sol at 49.5.
  • Tencent's deployment FAQ states that open-source release does not mean unrestricted commercial use and that preview status is not production approval.
The gist

Hy4 is among the largest open-weight models published, making a 770B-parameter MoE model available for download. Its coding benchmark results are competitive with several leading closed models, while graduate-level reasoning lags behind the top frontier tier.

What is moving now · Every edition · Every story

The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free