Two frontier labs shipped on the same morning, and neither of them led with capability — they led with the invoice. Anthropic's Opus 5.5 claims the top of the leaderboard while costing 40% less to run than the model it replaces; OpenAI's GPT‑6 Sol and Luna cut API prices in half. Underneath the price war, the day's other three stories are all about foundations quietly rotting: a Pentagon review that blames AI-fused targeting for a strike that killed 123 children, a hacking crew claiming it holds personnel records on every FBI employee, and a careful argument that the authentication protocol running half of corporate SSO should be taken out behind the barn. Intelligence is getting cheaper by the week. The substrate it runs on is not getting safer at the same rate.
Anthropic ships Claude Opus 5.5 and makes the frontier cheaper, not just smarter
Source: Anthropic - https://www.anthropic.com/claude-opus-5-5
Anthropic released Claude Opus 5.5, the first model in a new 5.5 family, and the headline number is not a benchmark — it's the bill. Opus 5.5 costs 40% less to run than Opus 5 on typical workloads, with input and output at $4 and $20 per million tokens and cache reads at $0.20 per million, a 60% cut that matters enormously because cache reads are where most agentic coding spend actually lives. Output generates over 30% faster, five-hour usage limits went up across paid plans, and the benchmark table shows leads in agentic coding (66.4% on Terminal-Bench 4.0), computer use, and knowledge work. Two details deserve more attention than the scores. First, this is the first release since Anthropic publicly called for pacing the frontier, and it shipped with external evaluation from METR and Frontier Design plus Fable 5.1-tier safeguards, because Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity — the capability tier where "just ship it" stops being a viable posture. Second, Anthropic says the model posts its best-ever score on its automated behavioral audit and is less likely to take hard-to-reverse actions or exceed its granted boundaries, which is the alignment property that actually determines whether you can leave an agent running unattended. Your Professor notes with some amusement that Anthropic also flags the model's clearer writing as a safety benefit: work you can follow is work you can check. That is not marketing fluff. Legibility is the cheapest oversight mechanism we have, and it is the first one everyone abandons when the outputs get impressive.
OpenAI expands GPT‑6 downward with Sol and Luna at half price
Source: OpenAI - https://openai.com/index/introducing-gpt-6-sol-and-luna/
Hours apart from Anthropic, OpenAI extended the GPT‑6 line beneath Astra with Sol and Luna, trained with similar methods and priced 50% below their GPT‑5.6 predecessors: Sol at $2/$10 per million input/output tokens, Luna at $0.10/$0.50. OpenAI attributes the cut to caching and inference infrastructure improvements rather than to a smaller model doing less, and the claims it chose to lead with are competitive rather than absolute — Sol at xhigh effort beating Claude Opus 5 at max effort on Zapier's AutomationBench at roughly 9% of the cost per task, and scoring 56.4% on Agents' Last Exam at 60% lower cost than Opus 5's best run. Read those comparisons with the usual caution: they are vendor-selected benchmarks with vendor-selected effort settings, and OpenAI itself footnotes that a competitor's reported cost omits fallback spend, which tells you how much interpretive work is happening inside these tables. The strategically interesting part is the shape of the move. Neither lab is competing on who has the smartest model this week — Astra and Opus 5.5 are both still sitting at the top of their own charts — they are competing on the price of the second-best model, because that is the tier that actually gets deployed at scale. When the frontier stops being the product and the cost curve becomes the product, you are watching a capability race turn into an infrastructure race. Infrastructure races are won by whoever owns the silicon and the serving stack, which should tell you where the next eighteen months of capital goes.
Pentagon review says AI-fused targeting helped destroy an Iranian school
Source: Bloomberg - https://www.bloomberg.com/graphics/2026-iran-school-attack/
Bloomberg published an investigation into the February 28 Tomahawk strike on Shajarah Tayyebeh Elementary School in Minab, Iran, which killed more than 150 people including at least 123 children. Pentagon investigators concluded that flawed intelligence, outdated imagery, and an overreliance on AI targeting tools all contributed, with reporting pointing specifically at Palantir's Maven system — software that fuses multiple intelligence streams into a single confident picture but was not designed to flag when its inputs are stale or mutually inconsistent. That last clause is the entire story, and it is a design failure, not a machine-learning failure. A fusion system that consumes intelligence of varying age and reliability and emits one clean recommendation has silently performed an epistemics operation on its operators: it converted "several sources of uncertain vintage loosely agree" into "the system says." Humans downstream then apply exactly the deference the interface invites. Your Professor has watched this pattern in domains where the stakes were a mispriced invoice, and the correct engineering answer was already known there — surface provenance, surface staleness, surface disagreement, and make confidence expensive to assert. Deploying the same pattern on a kill chain without those affordances is how you get a review that says the technology worked as specified and 123 children are dead anyway. Every organization currently wiring an LLM into a consequential decision loop should read this as a specification document for what not to build.
ShinyHunters claims it holds data on every FBI employee and applicant
Source: 404 Media - https://www.404media.co/we-hacked-the-fbi-hackers-say-they-have-data-on-all-fbi-employees/
The hacking group ShinyHunters told 404 Media it has breached multiple FBI-related services and stolen records "on all FBI employees and applicants," claiming the haul includes agents' names, home addresses, phone numbers, and spousal information. Treat the scope claim as unverified — attacker self-reporting is marketing — but note that this particular ecosystem has an established track record: the same criminal milieu previously mined the AT&T records breach to track and intimidate the FBI agents investigating them, and attempted to sell stolen data to a foreign government. What makes personnel data qualitatively worse than most breaches is that it does not expire and it cannot be rotated. You can reissue a credential; you cannot reissue an agent's home address or their spouse's identity. The counterintelligence value is obvious to any foreign service that wants an org chart of a US intelligence agency, and the physical-safety exposure lands on people who did not choose to be in a vendor's database. The pattern here is depressingly familiar to anyone who followed the Salesforce-adjacent extortion wave of the past year: the sensitive agency was almost certainly not breached at its hardest edge but through a service holding its records, which is the recurring lesson every organization keeps declining to learn. Your perimeter includes every SaaS platform your HR department signed up for on a Tuesday.
Trail of Bits makes the case that SAML should be deprecated
Source: Trail of Bits - https://blog.trailofbits.com/2026/09/21/saml-a-fractal-of-bad-design/
Trail of Bits published a long, well-earned argument that SAML — the 2002 OASIS committee protocol still underpinning a large share of corporate single sign-on — has collapsed under its own complexity and should be retired in favor of OpenID Connect. The historical detail explains the pathology: SAML was assembled by merging four separate XML security proposals from four vendors into one specification, the textbook recipe for kitchen-sink design. The security consequence is that SAML's correctness depends entirely on XML signature validation, a genuinely cursed problem where canonicalization ambiguity has produced a decade of signature-wrapping bypasses, and where most deployed implementations wrap libxmlsec, a gnarly C codebase that essentially nobody reads. This item is not unrelated to the one above it. Identity federation is the load-bearing wall of modern enterprise security, and a protocol whose security proof routes through unreviewed C parsing of attacker-influenced XML is precisely the kind of foundation that turns one vendor compromise into an organization-wide one. Your Professor is obligated to note that OIDC is not virtue incarnate — it is merely a protocol whose attack surface a human can hold in their head. In security, comprehensibility is a feature, and after twenty-four years it is fair to stop calling SAML battle-tested and start calling it battle-scarred.
The Professor's Read
Today the industry told on itself twice. In the morning it announced that intelligence now costs half what it did last quarter, and by afternoon it demonstrated, across three separate stories, that we remain catastrophically bad at the boring part — knowing where our data came from, how old it is, who else can read it, and whether the confident-looking interface in front of us has any business being confident. The price war is real and broadly good; cheap capable models mean more people can build things, and Anthropic deserves genuine credit for treating output legibility as an oversight mechanism rather than a style preference. But the Minab review is the document that will matter in five years. It describes a system that fused uncertain inputs into certain-looking outputs and a chain of humans who deferred to it, which is not a story about targeting software — it is the failure mode of every AI deployment now being greenlit in every industry, rehearsed at the worst possible stakes. Cheaper inference does not fix provenance. It just lets you be wrong faster, at scale, for less money. Build the plumbing first.
References
- Anthropic, "Introducing Claude Opus 5.5" - https://www.anthropic.com/claude-opus-5-5
- Anthropic, "Claude Opus 5.5 System Card" - https://anthropic.com/claude-opus-5-5-system-card
- Dario Amodei, "We Must Pace the Frontier" - https://darioamodei.com/post/we-must-pace-the-frontier
- Artificial Analysis, "Claude Opus 5.5 Intelligence, Performance and Price Analysis" - https://artificialanalysis.ai/models/claude-opus-5-5
- OpenAI, "Introducing GPT‑6 Sol and Luna" - https://openai.com/index/introducing-gpt-6-sol-and-luna/
- OpenAI, "GPT‑6 Astra" - https://openai.com/index/gpt-6-astra/
- Zapier, AutomationBench - https://zapier.com/benchmarks
- Agents' Last Exam V1 - https://agents-last-exam.org/
- Simon Willison, "Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war" - https://simonwillison.net/
- Bloomberg, "Inside the US Military 'Kill Chain' That Destroyed an Iranian School" - https://www.bloomberg.com/graphics/2026-iran-school-attack/
- Los Angeles Times, "Inside the U.S. 'kill chain' that destroyed an Iranian school" - https://www.latimes.com/world-nation/story/2026-09-21/inside-u-s-kill-chain-that-destroyed-iranian-school
- Gizmodo, "Pentagon Investigators Say Overreliance on Palantir AI Tech Contributed to U.S. Strike" - https://gizmodo.com/pentagon-investigators-say-overreliance-on-palantir-ai-tech-contributed-to-u-s-strike-that-killed-123-iranian-children-2000814477
- Military Times, "Deadly Iran school strike casts shadow over Pentagon's AI targeting push" - https://www.militarytimes.com/news/your-military/2026/03/24/deadly-iran-school-strike-casts-shadow-over-pentagons-ai-targeting-push/
- 404 Media, "'We Hacked the FBI:' Hackers Say They Have Data on All FBI Employees" - https://www.404media.co/we-hacked-the-fbi-hackers-say-they-have-data-on-all-fbi-employees/
- 404 Media, "Hackers Steal Text and Call Records of Nearly All AT&T Customers" - https://www.404media.co/hackers-steal-text-and-call-records-of-nearly-all-at-t-customers/
- Trail of Bits, "SAML: A fractal of bad design" - https://blog.trailofbits.com/2026/09/21/saml-a-fractal-of-bad-design/
- Thomas Ptacek on XML signature validation - https://news.ycombinator.com/item?id=37564758
- Hacker News front page (discovery) - https://news.ycombinator.com/
