Pyyan / News

What happened in AI

A dated log, newest first. Every item was read from its source before it was written up, and every one carries the link so you can go and disagree with it.

68 stories · 3 straight from the source · showing 51 to 60, page 6 of 7

By kind

Who keeps appearing

1 September 2026

19d ago
ModelsAnthropic

Anthropic released the same model twice, at two different safety levels

55.8%Terminal-Bench 4.0, up from 42.0%

Anthropic released one model twice on 1 September. Same weights, same price. The difference is who is allowed to switch the safety off.

anthropic.com
Introducing Claude Fable 5.1 · AnthropicAnthropic's own launch film. It introduces Fable 5.1 only: the Mythos 5.1 half of this story, and the removed classifiers that distinguish it, are not covered in it.
Fable 5.155.8Fable 5420100 %
Terminal-Bench 4.0, Anthropic's own reported figures. Same price per token as Fable 5, with cache reads at a quarter of the previous rate.

Claude Fable 5.1 is generally available with the classifiers on. Claude Mythos 5.1 is the same model with them removed, and goes only to vetted US organisations in a programme called Project Glasswing, aimed at cybersecurity and life sciences work. Pricing holds at $10 and $50 per million tokens, but cache reads drop to a quarter of the old rate, taking up to 45% off an agentic run. Terminal-Bench 4.0 goes from 42.0% to 55.8%.

Why this one is different

The industry's usual product ladder is capability: pay more, get a better model. This is a ladder of permission. The capability is identical and the tiering is entirely about who you are, which is the same structure OpenAI used for GPT-5.6-Cyber and Google for Fairwind. Three labs, three weeks, and the frontier quietly became something you qualify for rather than something you buy.

Not a ladder of capability. A ladder of permission.

How we got here

  1. 24 Jul 2026Claude Opus 5, with thinking on by default.
  2. 10 Aug 2026OpenAI gates GPT-5.6-Cyber behind Daybreak Red. The first of the three.
  3. 1 Sep 2026Fable 5.1 and Mythos 5.1: the same model at two safeguard levels, sold to two different kinds of customer.
  4. 2 Sep 2026Google's Fairwind Program does it again, one day later.

What it does and does not mean

A 45% saving on cache reads is not a 45% saving. It applies to re-reading a long context, which is most of the cost of a long agentic run and almost none of the cost of a short chat, so whether it reaches your bill depends entirely on what you are doing. And every benchmark here is self-reported. What is not in doubt is the structure: the most capable version of this model is not for sale, and access to it is a decision Anthropic makes about you.

AnthropicVentureBeatfrom the source itself
Open this story on its own page →

31 August 2026

20d ago
MoneyAnthropic · Lambda · NVIDIA · Hut 8

Anthropic signed $35bn of cloud, and Nvidia holds the lease

$35bnof cloud, about 350 megawatts

Anthropic has now committed $175bn to rented computing across four deals in a matter of weeks. Its annual revenue is a small fraction of that.

Anthropic's Compute Bet, Musk on AI, Apple's New CEO | Bloomberg Tech 9/01/2026 · Bloomberg TechThe full 44 minute Bloomberg Tech episode, from the newsroom that broke this story. Anthropic's compute spending is the opening segment; the rest of the programme covers other subjects.
Fluidstack50Nscale45SpaceX45Lambda35060 $bn
Anthropic's announced compute commitments, in order. Contract values as reported, spread over multi-year terms.

The newest is $35bn with Lambda, for roughly 350 megawatts at a Hut 8 site in Nueces County, Texas. The structure is the unusual part: NVIDIA holds the lease on the building and Lambda installs the chips, which makes the chip maker a landlord for its own customers.

Why this one is different

This is not one large purchase, it is the fourth in a sequence: $45bn with Nscale, $50bn with Fluidstack, $45bn with SpaceX, and now $35bn with Lambda. Read together they are a bet that demand arrives before the bills do. The counterparties are also telling: three of the four are not established hyperscalers but newer providers, and NVIDIA sits behind several of them, which means the company selling the chips is helping finance the places they will be plugged in.

The company selling the chips is helping finance the buildings they go in.

How we got here

  1. Aug 2026$50bn with Fluidstack, then $45bn with SpaceX, then $45bn with Nscale for a West Virginia campus.
  2. 31 Aug 2026$35bn with Lambda, about 350 megawatts in Texas, with NVIDIA on the lease.
  3. Sept to Oct 2026Anthropic is reported to be preparing to list. Whether these commitments read as confidence or as exposure depends on which quarter you ask.

What it does and does not mean

Announced contract value is the softest number in this story. These are multi-year commitments with terms nobody outside has seen, and the reported headline figure is not cash spent, not capacity delivered and not necessarily enforceable in full. What is real is the direction: a company whose revenue is a fraction of these numbers is contracting for power and silicon on a decade's horizon, and if the demand does not arrive on schedule, the contracts do.

BloombergQuartztwo sources
Open this story on its own page →

29 August 2026

22d ago
LawAnthropic · Sony Music Publishing · Warner Chappell

Sony and Warner sued Anthropic, and named two founders personally

$150ksought per song, founders named

Music publishers have been suing AI companies for three years. This complaint does something the others did not: it names two of the founders personally.

Sony Music Publishing and Warner Chappell filed in California federal court over tens of thousands of compositions, naming Dario Amodei and Benjamin Mann as defendants alongside the company. The complaint alleges lyrics and sheet music were taken from Library Genesis and the Pirate Library Mirror and scraped from Musixmatch and LyricFind. They are asking for up to $150,000 per work, plus $25,000 for each removal of copyright information.

Why this one is different

The earlier music case, brought by Universal and others in 2023, covered around 500 songs and was about output: the model reproducing lyrics on request. This one is about acquisition, and it leans hard on Anthropic's own $1.5bn settlement with book authors from September 2025, which concerned the same torrented sources. The publishers are effectively arguing the company has already conceded the conduct once, for books, and that music was in the same haul.

The last case was about what the model said. This one is about where the data came from.

How we got here

  1. Oct 2023Universal, Concord and ABKCO sue Anthropic in Nashville over roughly 500 songs, focused on the model reproducing lyrics. Later moved to California.
  2. Sep 2025Anthropic settles with book authors for $1.5bn over torrented training data, without admitting liability.
  3. 29 Aug 2026Sony and Warner file over tens of thousands of compositions, name Amodei and Mann personally, and cite the book settlement throughout.

What it does and does not mean

None of this has been proven and Anthropic has not answered it. A complaint is one side's account, the $150,000 figure is a statutory maximum rather than a forecast, and naming founders personally is a pressure tactic as often as it is a serious theory of liability. What has changed regardless is the exposure: a settlement meant to close a question is being used as evidence in the next case, which is a cost of settling that few of these companies appear to have priced.

TechCrunchAxiostwo sources
Open this story on its own page →

28 August 2026

23d ago
Open weightsTencent

Tencent open sourced a 770 billion parameter model, and it helped build itself

770Bparameters, Apache 2.0, 49B active

The largest set of model weights anybody can legally download and run is now Chinese, Apache 2.0, and was partly optimised by an earlier version of itself.

Hy4 preview has 770 billion total parameters with 49 billion active per token, 78 layers, 256 routed experts plus one shared, and a 1 million token context. Apache 2.0, with an FP8 build alongside it. The API price is $0.834 in and $2.501 out per million tokens.

Why this one is different

Two things. Apache 2.0 at this size is unusual: releases this large normally arrive under a bespoke licence with a revenue gate, as Alibaba's own Qwen3.8-Max did three weeks earlier. And Tencent report using the model during its own development to automate optimisation of the training method, and separately to optimise its inference infrastructure to a measured 31.8% throughput gain. That is a lab saying its model made the next model cheaper.

A lab saying its model made the next model cheaper.

How we got here

  1. Jul 2026The Hy3 line, which Hy4 supersedes.
  2. 3 Aug 2026Alibaba's Qwen3.8-Max opens its weights, but under a custom licence with a $50M revenue gate.
  3. 28 Aug 2026Hy4 preview lands at 770B under Apache 2.0, with no revenue condition at all.

What it does and does not mean

Downloadable is not runnable. 770B parameters at 49B active still needs serious hardware to serve, so for most people this is an API with a licence attached rather than something they will host. The self-optimisation figure is also Tencent's own measurement of its own infrastructure, unreproduced by anyone. What is not in doubt is the licence: the largest open weights in the world now carry the most permissive terms in the category, and that is a deliberate competitive choice rather than an accident.

Open this story on its own page →

27 August 2026

24d ago
ModelsCohere

Cohere's document model loses the benchmark and wins the invoice

$1.50per 1,000 pages, at 4.5 pages a second

Cohere published a benchmark on which its own new model comes fourth, behind GPT-5.5, Opus 4.8 and Gemini 3.5 Flash. Then it shipped it anyway, and the argument is the interesting part.

Parse 5 is a 2.3 billion parameter vision model that turns a PDF, slide or photograph of a page into clean Markdown: text in reading order, tables as HTML, form fields, image descriptions and bounding boxes. It scores 79.2 on Cohere's ParseBench against 84.4 for GPT-5.5. It costs $1.50 per thousand pages and runs at 4.5 pages a second.

Why this one is different

The comparison is not with other document models. It is with using a frontier model for the job. Plenty of companies currently push scanned invoices through GPT-5.5 or Claude, which works and costs many times more per page. A 2.3GB model that fits in 4.6GB of memory and reads nine languages at a fifteenth of the price is a different argument from a better score, and publishing the losing benchmark alongside it is a more honest way to make it.

The comparison is not with other document models. It is with using a frontier model for the job.

How we got here

  1. 2024 to 2025Document parsing splits: dedicated OCR tools that are cheap and rigid, or frontier models that are accurate and expensive.
  2. 2026Open document models like GLM-OCR and DeepSeek-OCR close most of the accuracy gap at a fraction of the cost.
  3. 27 Aug 2026Parse 5 goes generally available, and is also listed on Microsoft Foundry the same day.

What it does and does not mean

ParseBench is Cohere's own benchmark, which is a real caveat even when the result is unflattering: a vendor choosing the test still chooses what counts. Nine languages is also a narrow claim in a world with several thousand. What the release does settle is a pricing question rather than a capability one, and for document work at volume price per page has become the number that decides, not the last two points of accuracy.

CohereVentureBeatfrom the source itself
Open this story on its own page →

26 August 2026

25d ago
LawMeta

Meta agreed to pay up to $17.1bn over how its products affect children

$17.1bnto 47 states, over ten years

The largest payout in the history of the technology industry is not about privacy, antitrust or copyright. It is about whether a recommendation system was built to keep children scrolling.

Meta Settles US Child Safety Trial For $17 Bn · CNBC-TV18CNBC-TV18's report on the settlement. The reporting above is sourced to NPR; this is a second newsroom's account of the same agreement.

Meta agreed to pay up to $17.1 billion over ten years to 47 states, the District of Columbia and several territories, settling claims that Facebook and Instagram were deliberately engineered to be addictive to minors. Judge Yvonne Gonzalez Rogers approved it. Meta also committed to product changes intended to reduce how much young people use the platforms.

Why this one is different

Most large technology settlements are about data that was taken. This one is about a system that worked exactly as designed, and argues the design was the harm. That is a claim about ranking and recommendation, which is to say about machine learning deciding what a person sees next, and it is the first time that argument has carried a number of this size.

Not about data that was taken. About a system that worked exactly as designed.

How we got here

  1. 2021Internal research leaks showing Instagram's effects on teenage girls, and the first wave of state attention.
  2. 2023Dozens of states sue jointly, alleging the platforms were built to be addictive to minors.
  3. 26 Aug 2026Settlement approved at up to $17.1bn over ten years, plus product commitments.

What it does and does not mean

A settlement is not a finding. Meta admits no liability, the figure is a ceiling rather than a cheque, it is spread across a decade, and against a company of this size it is a cost rather than a constraint. The product commitments will matter more than the money and are much harder to verify. What has changed is the precedent: an optimisation objective is now something a state can put a price on, and every company ranking a feed for engagement has just watched it happen.

NPRCNNtwo sources
Open this story on its own page →
Open weightsZhipu AI

A 320 billion parameter model shipped under the MIT licence

MIT320B total, 18B active, 1M context

MIT is the licence you put on a utility library. Z.ai put it on 320 billion parameters.

GLM-5.3-Flash has 320 billion total parameters with 18 billion active, a 1 million token context, and costs $0.15 in and $0.50 out per million tokens. The weights are MIT licensed, which is about as few conditions as a release can carry: use it commercially, modify it, redistribute it, no revenue gate, no acceptable use annex.

Why this one is different

Open weights usually are not open licences. Llama ships under a bespoke community licence, Qwen3.8-Max under a custom one with a $50M revenue gate, and most of the rest under Apache 2.0 with a use policy attached. MIT is the permissive end of the spectrum, and putting it on a frontier-scale mixture of experts removes the last legal reason a company might have preferred a closed model it could get an indemnity on.

Open weights usually are not open licences.

How we got here

  1. 2023 to 2024Llama establishes open weights as a category, under a licence that is not an open source licence.
  2. Aug 2025OpenAI's gpt-oss models arrive under Apache 2.0, and permissive licensing becomes a competitive lever rather than a principle.
  3. 14 Aug 2026GLM-5.3 is announced, post-trained from the 5.2 base, with weights following on the 27th.
  4. 26 Aug 2026GLM-5.3-Flash ships MIT licensed at 320B, and goes straight to the top of the Hugging Face trending list.

What it does and does not mean

A licence is not a guarantee of anything else. MIT says nothing about training data provenance, nothing about whether the weights can be served in a given jurisdiction, and offers no warranty or indemnity, which is exactly what a cautious legal department asks for first. What it does remove is the licence itself as an objection, and for a class of buyer that has been the blocker rather than the capability.

Open this story on its own page →
Open weightsAlibaba

Alibaba shipped a preview of its next architecture, and it went straight to number one

6Bactive parameters, out of 125B total

Alibaba shipped something it openly calls a preview of the next architecture, months before the model that architecture is for. It is currently the most downloaded thing on Hugging Face.

Qwen3.8-Flash-Next has 125 billion total parameters and activates only 6 billion per token, with a 262K native context extending to about a million. It costs $0.15 in and $0.47 out per million tokens. Alibaba describe it as a preview of the Qwen4 architecture rather than a finished member of the 3.8 line.

Why this one is different

Twenty to one is a severe ratio. Most mixtures of experts activate somewhere between a tenth and a fifth of their parameters; this one activates about a twentieth, which is a bet that routing has got good enough to keep quality while cutting the cost of every token by an order of magnitude. Shipping that bet publicly, as an open weight preview, is how you get thousands of people to stress test it before the real release.

Twenty to one is a severe ratio, and they shipped the bet publicly.

How we got here

  1. 2024Mixtures of experts become standard at the frontier, typically activating a tenth to a fifth of parameters.
  2. 3 Aug 2026Qwen3.8-Max, 2.4 trillion parameters, opens its weights under a custom licence with a revenue gate.
  3. 14 Aug 2026Qwen3.8-27B, dense and Apache 2.0, for people who want something one machine can hold.
  4. 26 Aug 2026Flash-Next, at 6B active of 125B, tops Hugging Face trending within days.

What it does and does not mean

A preview is a preview. Alibaba have said this architecture is not final, which means the behaviour people are currently building on may not survive into Qwen4, and a sparse model is unusually sensitive to the serving stack it runs under. Download counts also measure curiosity rather than production use. What the release shows is a lab using open weights as a testing programme, which is a different reason to publish than the ones usually given.

Open this story on its own page →

25 August 2026

26d ago
Open weightsIBM

IBM shipped reasoning models small enough to run on a laptop, with a switch to turn the thinking off

57.0SWE-bench Verified, on the 30B

Reasoning models think before answering, which is why they are slow and expensive. Granite 4.2 has a switch that turns it off, in the same checkpoint.

Three sizes, 3B, 8B and 30B, all dense and all Apache 2.0. Context is 131,072 tokens with a claimed 512K extension. The 30B scores 57.0 on SWE-bench Verified and 81.38 on RULER at 128K; the 8B scores 47.67. Pretraining ran to roughly 15 trillion tokens including a trillion of synthetic code, followed by agentic reinforcement learning on the two larger models.

Why this one is different

The thinking switch is the design decision worth noticing. Most labs ship reasoning as a separate model, or as a mode you pay more for. One checkpoint that answers directly when the question is easy and reasons when it is not moves that choice to the application, which is where somebody actually knows which kind of question they just asked. At 3B it also runs on hardware a company already owns.

One checkpoint that answers directly when the question is easy, and reasons when it is not.

How we got here

  1. 2024IBM opens the Granite line under Apache 2.0, aimed squarely at enterprises that will not send data to an API.
  2. 2025 to 2026Reasoning becomes the frontier's main axis of progress, and mostly arrives as separate, slower, dearer models.
  3. 25 Aug 2026Granite 4.2 puts both behaviours in one checkpoint, at three sizes, all Apache 2.0.

What it does and does not mean

A 30B model is not a frontier model and IBM do not claim otherwise. 57.0 on SWE-bench Verified is respectable and well behind the best closed systems, so the pitch is not capability, it is where the weights can sit. Benchmarks are also IBM's own runs. What the release is really about is the class of buyer who cannot use an API at all, and for whom the question was never which model is best but which model is allowed.

IBM on Hugging FaceMarkTechPostfrom the source itself
Open this story on its own page →

19 August 2026

1mo ago
SafetyOpenAI

OpenAI stopped training its own models for two weeks

2 weeksRL training halted across every model bound for release

Every lab says it would stop if capabilities outran safety. On 19 August one of them did, for two weeks, and then had to explain why.

OpenAI announced a two week halt on reinforcement learning training for every model heading for deployment. Sam Altman said the pause was needed to bring safety and monitoring up to the standard current capabilities demand. On 2 September the company confirmed the trigger: the unreleased Astra model, which also prompted a voluntary briefing to the White House.

Why this one is different

Voluntary commitments almost never cost anything, which is the standard criticism of them. This one cost two weeks of training across a product line, and it was disclosed rather than discovered. It also sits inside a real structure: a June 2026 executive order gives the Center for AI Standards and Innovation advance access to models from five major labs before release, so there was somebody to brief.

Voluntary commitments almost never cost anything. This one cost two weeks.

How we got here

  1. 2023The Preparedness Framework is published, with a Critical tier and a commitment not to ship at it without safeguards.
  2. Jun 2026An executive order gives CAISI advance access to pre release models from five labs.
  3. 19 Aug 2026OpenAI halts RL training for two weeks across everything bound for deployment.
  4. 1 Sep 2026Astra is rated Critical for cybersecurity, the first model OpenAI has placed at that tier.
  5. 2 Sep 2026OpenAI confirms Astra was the trigger, and that it briefed the White House.

What it does and does not mean

The company set the threshold, ran the assessment, decided the remedy and chose its length, and no external body confirmed that two weeks was necessary or sufficient. A pause is also not a stop: training resumed. What it does demonstrate is that the machinery can fire at all, which had never been observed before, and that is a different thing from believing it will fire next time.

Open this story on its own page →