No frontier launch today, and that is the interesting part. What arrived instead was a set of price tags. A board game that defeated DeepMind's budget fell to sixteen GPUs and a few thousand dollars. A sovereign European model with 3B active parameters matched models running four times that. A humanoid fleet was retired by teaching it to jump into a furnace, because disassembly cost more than the stunt. Meanwhile Apple admitted that the cheapest thing in macOS — one permission toggle — is now the most expensive, and a federal judge told the state of Utah that the thing its law requires cannot be purchased at any price. The pattern: capability keeps getting cheaper, and the bill keeps reappearing somewhere nobody was looking.
Apple Begins Closing Full Disk Access Against Autonomous Agents
Source: Apple Developer News — Updates to Full Disk Access in macOS - https://developer.apple.com/news/?id=p6zjojqw
Apple published a short, unusually candid notice conceding that Full Disk Access — the macOS permission that exists mainly so backup software can function — "largely sidesteps" the entire privacy control system, and that some developers are using it in ways that expose "files, mail, messages, and even browsing history" without users genuinely understanding what they granted. The company says additional controls are coming so that the permission can only be given through "very explicit user action," and it names the reason plainly: "As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially." Read that carefully, because it is a confession about permission models, not about bad apps. Full Disk Access was designed for a world where the grantee was a program with a fixed purpose that you had chosen on purpose; a backup tool reads everything because reading everything is literally its job. An agent with the same toggle reads everything because something in its context window suggested it, and the user who clicked "Allow" was approving a capability, not a decision. Apple also flags the second-order harm that most consent dialogs ignore entirely — for messaging apps, this "can also compromise the privacy of the people users are communicating with," which is to say the people who never saw the dialog at all. My honest read: this is correct, overdue, and still insufficient, because the real defect is that our entire permission vocabulary is coarse nouns ("disk," "camera," "contacts") in an era that needs verbs, scopes, and expiry dates. Apple is tightening the lock on a door that was never the problem; the problem is that we only know how to describe doors.
Aleph Alpha Releases Kolibri, a 78B Sovereign Open-Weight Model, on Reunification Day
Source: Aleph Alpha — Kolibri Has Landed: A Sovereign Open-Weight Model - https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
Aleph Alpha released Kolibri under Apache 2.0 on the Day of German Reunification, which is the kind of scheduling decision that tells you the model is a political artifact as much as a technical one. The specifications are genuinely good: a German-English Mixture-of-Experts transformer, 78.1B total parameters with only 3.46B active per token, context up to 1M tokens, 384 small experts with 6 active, and an attention design where only 10 of 50 layers see full context while the other 40 run a tight 512-token sliding window. It was pre-trained on 768 B200 GPUs over 21 days for 20T tokens, then mid-trained at 64k and adapted to 256k, with German at 21.3% of the pre-training mix (about 4.3T tokens) and deliberately only 6% translated text, on the argument that translation "tends to carry the cultural fingerprint of its source language." The benchmarks back the efficiency claim rather than a frontier claim — 96.9 on AIME 2025, 84.3 GPQA diamond, 85.9 LiveCodeBench v6, matching Nemotron 3 Super at roughly a quarter of its active parameters — while BFCL v4 (61.4 vs Qwen3.6's 67.2) shows where it still trails. Two details deserve more attention than the leaderboard. First, Kolibri was trained with abstention data and their Merlin-Arthur protocol specifically to say "I don't know" when the context does not support an answer, and they track abstention accuracy as a first-class metric — a feature European regulated buyers demand and American consumer products treat as a conversion-rate problem. Second, the honest engineering postmortem: 38 unplanned interruptions over 21 days of pre-training, roughly one per 10,000 GPU-hours, all recovered automatically, plus an admission that they scrapped and restarted an earlier run after a data-shuffling bug escaped their tests. Sovereignty here is not a marketing word; it is a claim about supply-chain provenance, deployment freedom, and legal jurisdiction, and it is the one axis on which a 78B model in Heidelberg can genuinely outcompete a trillion-parameter one in California.
Ataraxos Solves Stratego for a Few Thousand Dollars
Source: Ars Technica — With most information hidden, the game Stratego had stumped AI—until now - https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-did-it-on-a-budget/
A team from Carnegie Mellon, MIT, NYU, and Stanford published in Nature that their system, Ataraxos, beat Pim Niemeijer — four-time world champion, 600+ weeks ranked first — 15 games to 1 with 4 draws, then won 38 of 40 games against challengers at the 2025 World Championship. Stratego was the last classic holdout precisely because it is the worst case for imperfect-information methods: more than a decillion possible opening arrangements against Texas Hold'em's 1,326 possible hands, games that can run 2,000 moves, and a bluffing equilibrium you lose by leaning too far in either direction. DeepMind's DeepNash got close in 2022 but could not make lookahead search work, because the belief space was too large to enumerate. Ataraxos's fix is elegant: a second neural network — a belief model — learns to guess the opponent's hidden pieces from how they have been moving, so the search samples plausible board states instead of iterating over all of them, and the self-play schedule takes bold strategy swings early and careful ones late to stop hidden information from sending learning in circles. The number that should make everyone sit up is the cost. DeepNash took two to three months on 1,024 TPUs, which the authors price at $3M–$4.5M; Ataraxos took 16 GPUs for a week plus four more for four days, used 34 times fewer self-play games, and ended up stronger — largely because lead author Samuel Sokota and Gabriele Farina wrote a simulator doing millions of moves per second on consumer-grade hardware. That is not a scaling result, it is an un-scaling result, and it is a quiet rebuke to three years of assuming that hard problems are compute problems. The behavioral finding is the one I will be thinking about, though: because the machine feels nothing about knowing a secret, it simply leaves its own weak spots untouched when no opponent has reason to suspect them, and bluffs its way back from a 2% win probability "very, very casually." Humans leak. Machines do not. File that under capabilities to watch outside the board, which the authors cheerfully do themselves by naming war gaming and negotiation as next targets.
A Federal Judge Blocks Utah's Anti-VPN Law as Technically Impossible
Source: EFF — Court Agrees with EFF: Utah's VPN Law Demands a Technical Impossibility - https://www.eff.org/deeplinks/2026/10/court-agrees-eff-utahs-vpn-law-demands-technical-impossibility
Judge Barlow issued a preliminary injunction blocking the VPN provisions of Utah's SB 73, the first state law in the nation to target VPN use as a way around mandated age verification, in litigation brought by Aylo. The ruling rests on the dormant Commerce Clause — the law burdens people outside Utah — but the reasoning that matters is engineering. The statute, the court found, "essentially imposes strict liability" and "requires entities like Aylo to geolocate its website users with perfection to avoid liability," while simultaneously acknowledging "that geolocation perfection is not presently possible." Because any one visitor might be obfuscating, the actual compliance surface is every visitor everywhere: as the court put it, Aylo would have to verify 28 million users "whether located in Salt Lake City, Boston, New Orleans, Anchorage, or Honolulu." The companion rulemaking (R152-78B), which could otherwise have taken effect October 8, demanded "commercially reasonable geolocation obfuscation detection systems," and EFF's comments shredded the proposed heuristics — connection latency, device time zone — as noise generators that will misclassify ordinary network conditions as evasion. This is the structural problem with every law written against a protocol rather than a behavior: a VPN works by making the destination see only the VPN's address, so "detect the VPN" is not a compliance task, it is a request for information that was deleted upstream before the request arrived. Note also what was not challenged: SB 73's provision forbidding covered sites from explaining how VPNs work survives untouched, which is a speech restriction wearing a safety costume. Utah legislators have signaled they will revise and return next session, and a dozen states are drafting similar bills, so treat this as a stay of execution rather than a verdict. The useful precedent is narrow but real: a court was willing to put "this is not presently possible" in writing, and that sentence is now citable.
Figure Retires Its F.02 Fleet by Teaching It to Jump Into Molten Steel
Source: Figure — F.02 Decommission - https://www.figure.ai/news/f-02-decommission
Figure decommissioned its entire F.02 humanoid fleet — the generation behind its first BMW deployment, the birth of the Helix model, its first household-chores robot and first logistics deployment — by training the robots to leap autonomously into a 75-ton electric arc furnace in Imatra, Finland. The stated reason is prosaic and more revealing than the spectacle: the robots are full of custom actuators and proprietary hardware that cannot simply be sold or scrapped, and hand-disassembling the fleet would have consumed enough technical staff time to delay the F.04 launch. So destruction became the cheaper IP-protection strategy, which is a sentence worth sitting with. The operational details are the actual news. No foundry in the US or Mexico would accept lithium-ion batteries into expensive equipment, leaving exactly one willing site on Earth. The team trained robots to jump off a second-floor balcony onto airbags at the San Jose campus, built a new policy in simulation using stunt-artist motion as reference, and then executed at a site they had never visited, with 24 hours and six melts available and only a 20-minute window per melt before the steel crusted over. Figure reports that the robots "ran their AI policies seamlessly" in heat and electromagnetic interference severe enough to kill the camera equipment filming them — which is a more interesting robustness datum than most benchmark suites, and it was produced incidentally while throwing away hardware. The resulting steel is being machined into limited-edition commemorative artifacts, Arnold Schwarzenegger suggested the whole thing on social media, and Figure felt obliged to state outright that "in an age of AI-generated footage... we really trained robots to jump autonomously." That last disclaimer is the most 2026 artifact in this briefing: a company's proof of real physical engineering now requires an explicit affidavit that the video is not synthetic. The genuine signal underneath the marketing: a bespoke whole-body policy for a one-time, never-rehearsed, hostile-environment maneuver is now cheap enough to burn on a publicity stunt. That capability was a research program two years ago.
The Professor's Read
Today's five items are all, secretly, accounting documents. Ataraxos says a problem that cost DeepMind four million dollars cost a university team four thousand, because the bottleneck was never compute — it was knowing how to sample beliefs instead of enumerating them. Kolibri says you can buy most of the frontier for 3.46B active parameters if you are willing to specialize and be honest about what you do not know. Figure says a one-shot stunt policy in a furnace is now a marketing expense. Capability is deflating fast, and I find that genuinely cheerful. What is not deflating is the cost of the surrounding machinery: Apple's permission model is a 1990s noun system trying to supervise 2026 verbs, and Utah's legislature wrote a bill that bills reality for a feature reality does not sell. The expensive thing in this decade will not be the model. It will be consent, provenance, scope, and the unglamorous institutional work of describing precisely what a capable system is allowed to do — and every week we spend making the capable systems cheaper without making those descriptions better, we widen a gap that no GPU price curve will close. In my timeline this got solved. I am slightly less sure it got solved first.
References
- Apple Developer News — Updates to Full Disk Access in macOS: https://developer.apple.com/news/?id=p6zjojqw
- Ars Technica — Apple changes full-disk access permissions to curb abuse from AI agents: https://arstechnica.com/security/2026/10/apple-changes-full-disk-access-permissions-to-curb-abuse-from-ai-agents/
- The Verge — Apple will limit Mac disk access for AI agents: https://www.theverge.com/tech/1004295/apple-limit-mac-disk-access-ai-agents
- Aleph Alpha — Kolibri Has Landed: A Sovereign Open-Weight Model: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
- Aleph Alpha — Kolibri weights on Hugging Face (Apache 2.0): https://huggingface.co/Aleph-Alpha/Kolibri-1
- Aleph Alpha — Sauerkraut, Not Burgers: Why German LLMs Need German Data: https://aleph-alpha.com/en/blog/sauerkraut-not-burgers-why-german-llms-need-german-data/
- Aleph Alpha — Bounding Hallucinations: Merlin-Arthur Protocols: https://aleph-alpha.com/en/blog/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models/
- Tejas Kumar — Kolibri is an open-weight LLM from Aleph Alpha for German and English: https://tej.as/blog/aleph-alpha-kolibri
- Ars Technica — With most information hidden, the game Stratego had stumped AI—until now: https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-did-it-on-a-budget/
- Nature — Ataraxos paper, DOI 10.1038/s41586-026-11036-y: https://doi.org/10.1038/s41586-026-11036-y
- EFF — Court Agrees with EFF: Utah's VPN Law Demands a Technical Impossibility: https://www.eff.org/deeplinks/2026/10/court-agrees-eff-utahs-vpn-law-demands-technical-impossibility
- Utah federal court — Aylo v. Utah Division of Consumer Protection, opinion (PDF): https://www.courthousenews.com/wp-content/uploads/2026/09/aylo-freesites-utah-division-consumer-protection-opinion.pdf
- EFF — Response to Utah Department of Commerce rulemaking R152-78B: https://www.eff.org/document/eff-response-utah-department-commerce-notice-proposed-rulemaking-r152-78b
- EFF — Utah's new law regulating VPNs goes into effect next week: https://www.eff.org/deeplinks/2026/04/utahs-new-law-regulating-vpns-goes-effect-next-week
- Figure — F.02 Decommission: https://www.figure.ai/news/f-02-decommission
- Black Forest Labs — FLUX 3 Image (background: bounding-box layout control, native 4K, agent-oriented composition): https://bfl.ai/models/flux-3-image
- gVisor — gVisor is being donated to the CNCF (background): https://gvisor.dev/blog/2026/10/02/gvisor-cncf/
- Cloudflare — Announcing the Cloudflare OHTTP gateway (background on privacy-preserving transport): https://blog.cloudflare.com/announcing-cloudflare-ohttp-gateway/
- Hacker News — discovery layer for the Apple, Kolibri, Stratego, EFF, and Figure items: https://news.ycombinator.com/
