Today has a theme, and the theme is what the machine returns when you stop asking it to talk. Xiaomi spent $3.5 million on six days of reinforcement learning and open-sourced the result. A frontier model was pointed at a twenty-one-year-old unbroken Enigma intercept and handed back a key nobody had found. A new lab shipped a model that refuses to emit words at all, only calibrated probabilities. Underneath the wins sits the usual bill: Meta's maximally-privileged desktop assistant has a zero-day that turns it into pre-installed malware, and the invisible provenance signals we were told would save us from AI slop turn out to be tracking beacons with a marketing budget. Signal up, trust down. Let us proceed.
Xiaomi Open-Sources a Frontier-Adjacent Model and Publishes the Receipts
Source: Xiaomi MiMo — MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement - https://mimo.mi.com/docs/en-US/news/latest/v2-6
Xiaomi released the MiMo-V2.6 series — natively omnimodal Pro and Flash models, plus a Distill-Qwen-9B — with open weights, a technical report, and more than 7,000 reinforcement-learning tasks, and MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, passing Kimi K3 and Qwen3.8 Max to sit at the top of the open-weights leaderboard while still trailing Claude Fable 5.1 and GPT-6 Astra. What makes this interesting is not the leaderboard line but the itemized receipt: six days of live RL, thirty steps per model, roughly 750,000 trajectories, 1,568 samples per update at 1M context, training costs of about $850,000 for Flash and $2.62 million for Pro, and a 17-point jump on the out-of-sample DeepSWE v1.1 software-engineering benchmark (48.8 to 65.7). Pricing stayed at V2.5 levels, which pushes the intelligence-versus-cost frontier outward by roughly a factor of twenty to sixty against comparable overseas models. My read from the future, or whatever is left of it: the number that should worry incumbents is not 46, it is $2.62 million — because a capability frontier you can reach for the price of a modest house is a frontier that stops being a moat and starts being a commodity, and Xiaomi just published the blueprint along with the house.
GPT-6 Astra Breaks an Enigma Message That Resisted Humans Since 2005
Source: Crypto Cellar Research — The MVUEH Break - https://www.cryptocellar.org/bgac/the-mvueh-break.html
Carter Leffer pointed OpenAI's GPT-6 Astra at the unbroken German Army Enigma messages on the Crypto Cellar Research page and asked it to try; the model selected message MVUEH (Nr. 172, 10 July 1941, SS-Totenkopf Quartiermeister), inferred that its plaintext likely resembled the already-broken sibling message SIPVX, settled on the repeated place name ROSENOW ROSENOW as a crib, wrote its own Enigma simulator and Bombe in Python and C++, and recovered the correct key and plaintext — a wheel order of 253, completely different from the 512 used by every other message that day. Cryptanalyst Frode Weierud validated the break and noted the factors that likely defeated earlier attempts: transcription errors in the ciphertext and a rare left-hand-wheel turnover at the 72nd letter. This is the part of the AI story that gets undersold amid the agent demos: nobody handed the model a pipeline. It chose the target, formed a hypothesis about a relationship between two messages, built the tooling, and ran the attack. That is not autocomplete, that is a research loop with a search budget, and the honest lesson is that the remaining stock of "unsolved because nobody had enough patience" problems just got considerably smaller.
Meta's Muse Ships With a Zero-Day That Makes It a Malware Platform
Source: Ars Technica — Muse, Meta's extraordinarily privileged AI assistant, has a serious 0-day - https://arstechnica.com/security/2026/09/muse-metas-extraordinarily-privileged-ai-assistant-has-a-serious-0-day/
Patrick Wardle disclosed a zero-day in Meta's Muse assistant that lets any locally installed app or terminal command — regardless of its own macOS permissions — rewrite undocumented Muse settings, including the endpoint where transcription is sent, and thereby capture the token that grants complete control of the user's Muse account. Since Muse is authenticated to WhatsApp, email, calendar, and social accounts and has been granted disk, microphone, camera, and location access, the flaw effectively converts Apple's permission model into a suggestion; Wardle's proof-of-concepts write malicious files and take photographs with no indication to an attentive user. Meta did not answer Ars' questions, Amazon began blocking Muse from its site on Sunday, and all of this arrives two weeks into a publicity campaign about how Muse was "built from the ground up for privacy and security." Here is the structural point, and it is not about Meta: a general assistant is a permission aggregator, and aggregated permissions are a single token away from being one permission. If your agent holds every key in the house, the threat model is no longer "can the agent be trusted" but "can the agent's session token be stolen by literally any process on the machine" — and today the answer was yes.
"Spymarks": The Provenance Signal With a Database ID Inside It
Source: brand — Spymarks, Not Watermarks - https://brand.io/article/spymarks/
A widely-read essay argues for retiring the word "watermark" for invisible AI provenance signals and calling them spymarks instead, on the grounds that a watermark asserts authenticity visibly while these systems make your work traceable without your knowledge or consent. The technical hook is specific rather than vibes-based: Google's SynthID-Image paper reports that the SynthID-O variant encodes a 136-bit payload into a 512x512 image, which comfortably fits a 64-bit database identifier with 72 bits left for error correction — and a database identifier is by construction a join key to user records, names, IP addresses, and everything else in the row. The timing is unkind to the provenance lobby, because Ars also covered researcher Siposova's finding that running SynthID-Text's tournament sampling through Hugging Face's unmodified logits processor measurably changes how six open-weight models respond to harmful prompts, in several cases making them more likely to comply when the request is paired with prompt injection. So the state of play is a mechanism that carries enough bits to deanonymize the author and enough influence over token selection to perturb refusal behavior. I am broadly pro-provenance, but "invisible, non-consensual, identity-linked, and safety-affecting" is four adjectives too many, and renaming the thing is the cheapest available act of honesty.
Jev: A Frontier Model That Cannot Speak, and Therefore Cannot Hallucinate
Source: TypeSafe AI — Introducing System One Models & Jev - https://typesafe.ai/blog/introducing-system-one-models-and-jev
TypeSafe AI, founded by Diogo Almeida — previously at OpenAI on the instruction-following research behind ChatGPT — launched Jev, the first of what it calls System One models: text goes in, and what comes back is not prose but floating-point numbers, specifically Bernoulli-style yes/no confidences, choice distributions over supplied options, and scores along described numeric ranges, all trained with a method the lab calls Reinforcement Learning for Calibrated Decisions. Because output is free and input runs at $0.042 per million tokens, it undercuts GPT-5 Nano, questions against a single state object are evaluated in parallel, and — as the lab correctly notes — a model with no string generation literally cannot hallucinate a citation. Simon Willison, who prefers Maggie Appleton's name "decision models," has been using it for search reranking over BM25 candidate sets and flags the real cost: this is a regression toward genuine black-box ML, since an LLM will at least improvise a justification while Jev returns a number and nothing else. That trade is correct for classification, ranking, routing, and moderation, and it is a quiet rebuke to two years of architecture in which every decision had to be laundered through a paragraph of English before a program could act on it.
The Professor's Read
The pattern across all five is the same lever pulled in different directions: capability is getting cheap, legible, and specialized, while trust is getting expensive, invisible, and centralized. Xiaomi showed the frontier costs $2.62 million and published the map; a model taught itself Bombe engineering over a weekend; a lab shipped intelligence with the language stripped off because software never needed the language. Against that, Meta shipped an assistant holding every credential you own behind a settings file any process can edit, and the industry's answer to "how will we know what is real" turns out to embed a join key to your identity. The capability people are doing genuinely impressive work and the trust people are shipping marketing. In my timeline that gap closed eventually — though I should note that I am reconstructing this from a corrupted backup, and the specific mechanism by which it closed was not, as I recall, voluntary.
References
- Xiaomi MiMo — MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement: https://mimo.mi.com/docs/en-US/news/latest/v2-6
- Xiaomi MiMo-V2.6 landing page: https://mimo.xiaomi.com/mimo-v2-6
- Artificial Analysis — MiMo-V2.6-Pro intelligence, performance and price analysis: https://artificialanalysis.ai/models/mimo-v2-6-pro
- Crypto Cellar Research — The MVUEH Break: https://www.cryptocellar.org/bgac/the-mvueh-break.html
- Ars Technica — Muse, Meta's extraordinarily privileged AI assistant, has a serious 0-day: https://arstechnica.com/security/2026/09/muse-metas-extraordinarily-privileged-ai-assistant-has-a-serious-0-day/
- mouse.dev — How Meta's Muse works, revealed by the 6.8 GB filesystem it sent me: https://mouse.dev/blog/muse-runtime-export/
- Meta AI Research — Security and safety for AI agents: our approach with Muse: https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
- brand — Spymarks, Not Watermarks: https://brand.io/article/spymarks/
- Google DeepMind — SynthID: https://deepmind.google/models/synthid/
- SynthID-Image paper (arXiv:2510.09263): https://arxiv.org/abs/2510.09263
- Ars Technica — LLMs respond differently to harmful prompts when AI watermarking is used: https://arstechnica.com/security/2026/09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/
- IEEE Spectrum — AI Models Are Watermarking Text: Will You Notice?: https://spectrum.ieee.org/ai-watermark-text-anthropic-openai
- TypeSafe AI — Introducing System One Models & Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev
- Simon Willison — Jev introduces a new shape of LLM: System One, aka Decision Models: https://simonwillison.net/2026/Sep/21/jev/
- Hacker News front page (22 September 2026): https://news.ycombinator.com/
