Latent.Space

Latent.Space

AINews: Weekday Roundups

[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%

overshadowing more efficient GPT6 models from OpenAI

Sep 23, 2026
∙ Paid

OpenAI made a valiant effort with GPT-6 Sol and Luna launching 50% lower than GPT-5.6, but with 17M views on the launch and counting, today was always going to belong to Claude Opus 5.5, “the first model in our new Claude 5.5 family” performing like “Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.”

X avatar for @claudeai
Claude@claudeai
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
4:31 PM · Sep 22, 2026 · 17.2M Views

2.74K Replies · 8.01K Reposts · 85.2K Likes

Opus 5.5 beats Fable or challenges Astra at most benchmarks, and both labs credited efficiency work for the API price cuts, but there are HUGE double digit gains everywhere from prefill to decode to overall compute…

X avatar for @theo
Theo - t3.gg@theo
Opus 5.5 is SMALLER than Opus 5?? Did Anthropic massively level up their post training? Huge.
6:53 PM · Sep 22, 2026 · 110K Views

127 Replies · 74 Reposts · 3.88K Likes

… with offsetting inefficiency in token usage on some frontier tasks.

X avatar for @ArtificialAnlys
Artificial Analysis@ArtificialAnlys
Claude Opus 5.5 (max) costs $5.98 per Intelligence Index task, which is similar to Opus 5 (max) at $5.86, but this bundles a significant token usage increase with Anthropic’s price reductions Compared to Opus 5, Opus 5.5’s increased token usage would drive an ~80% increase in …
11:34 PM · Sep 22, 2026 · 36.1K Views

39 Replies · 25 Reposts · 603 Likes

HOWEVER something that is a rare emphasis in the Claude launch was the writing improvements: “It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.”

We can confirm - here is today’s AINews section run on Opus 5.5 and Sol 6. The difference is night and day - we are migrating to Opus 5.5 immediately for AINews going forward until we reach the next model/version of AINews.

They have also published initial work on large multiagent swarms (and efficiency):

X avatar for @scaling01
Lisan al Gaib@scaling01
very proud of Anthropic bros to be the first lab to report multi-agent scaling up to 100 parallel agents in their system card
X avatar for @scaling01
Lisan al Gaib @scaling01
Opus 5.5 System Card https://t.co/D1PDkVqChC
5:01 PM · Sep 22, 2026 · 82.2K Views

30 Replies · 60 Reposts · 1.19K Likes

AI News for 9/21/2026-9/22/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!


AI Twitter Recap

Top Story: Claude Opus 5.5 launch, numbers, and reactions

What happened

Anthropic shipped Claude Opus 5.5, the first model in a new Claude 5.5 family. Its pitch is Fable 5.1‑level capability at Opus pricing, with more speed and better writing. OpenAI released GPT‑6 Sol and Luna about an hour later.

  • Launch claims. Opus 5.5 “performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5” (@claudeai; @AnthropicAI).

  • Where it leads. Anthropic says it leads on agentic coding, computer use, and knowledge work (@claudeai).

  • Speed and cost. It is about 30% faster and about 40% cheaper per task than Opus 5 (@ClaudeDevs, @lydiahallie).

  • Communication fixes. The model puts the most important information up front and follows user writing rules. This targets the most common feedback on Opus 5 (@claudeai).

  • Subscription changes:

    • 5‑hour session limits are up 20%.

    • Lower pricing means limits go 25% further.

    • Pro, Max, and Team users get a banked rate‑limit reset they can use whenever they choose (@claudeai, @ClaudeDevs, @trq212).

  • New defaults. Opus 5.5 is now the default in Claude Code and the Claude app, including Cowork. Default effort is medium, described as “comparable to Fable 5.1 on intelligence but faster” (@_catwu).

  • Availability. It is live in Claude Code and the Claude Platform API (@ClaudeDevs), and in Claude Tag for Slack (@_catwu).

  • Roadmap. Sonnet 5.5 and Haiku 5.5 follow “in the coming weeks” (@mikeyk, @AiBattle_). This contradicts rumors that Haiku was discontinued (@kimmonismus).

  • Safeguards. Opus 5.5 is the first Opus with Fable 5.1‑class safeguards on cyber, bio, and frontier LLM development. Flagged requests fall back to another model, and Anthropic says it is “working to reduce incorrect flags” (@ClaudeDevs).

  • Pre-release signals. The model was spotted in Claude Code shortly before the announcement (@kimmonismus).

  • System card. It was published at launch (@scaling01).

Pricing and token economics (facts)

  • List price. Token pricing was cut 20%, from $5/$25 to $4/$20 per 1M input/output tokens (@ValsAI).

  • Offset by higher token use. Vals notes Opus 5.5 often uses more tokens, especially on coding, where it posts its largest gains. The lower sticker price is partly offset by usage.

  • Artificial Analysis cost breakdown. At max effort, Opus 5.5 costs $5.98 per Intelligence Index task versus $5.86 for Opus 5 (max). Their decomposition (@ArtificialAnlys):

    • Higher token usage alone would raise cost per task about 80%, to $10.51.

    • The 20% base-price cut brings that to $8.41.

    • Cheaper cache reads ($0.20) bring it to $5.98.

  • What that means. At max effort, the per‑task saving over Opus 5 disappears. The “40% cheaper” claim applies to default (medium) settings.

  • Relative to Fable 5.1. Cline reports Opus 5.5 beats Fable 5.1 on the Artificial Analysis Intelligence Index at about 2.5x lower cost (@cline).

  • Prompt caching. Switching effort mid‑session does not break the prompt cache on Claude Code v2.1.280+ (@lydiahallie).

  • Model size (speculation). @theo claimed Opus 5.5 is smaller than Opus 5 and credited post‑training. This was not confirmed in official posts.

Benchmarks and independent evals

Anthropic’s own table. Opus 5.5 beats Fable 5.1 on every row of Anthropic’s headline comparison and beats GPT‑6 Astra on most (@kimmonismus, @synthwavedd, @scaling01).

@ShayneRedford (Anthropic) summarized the claimed gains:

  • Stronger than Astra on CursorBench, KWBench, and OSWorld.

  • Much better style and instruction following.

  • Stronger science and health capabilities.

  • More robust against cyber and bio misuse.

Third‑party and partner evals:

EvalResultSourceVals Index#1, up 2 spots / 2 pts vs Opus 5; Anthropic holds the top three spots (GPT‑6 Sol pending)@ValsAIVals RSI Index#1; first model to beat the published reference on LM Training under their protocol; beats Fable 5.1@ValsAI, @ValsAIFrontierSWE (Proximal)62.3%, #2 behind GPT‑6 Astra (65.5%); ahead of Fable 5.1 (56.3%) and Opus 5 (52.0%)@ProximalHQFrontierCode 1.1 (Cognition)65.3% on Extended; takes #1 from Fable 5 “at a fraction of the cost”@cognitionCursorBench57.8% (Max), new top model; 40% less per task than Opus 5@cursor_aiPerplexity WANDR0.610 at $4.13/task; slightly above Fable 5.1 at 67.6% lower cost@perplexity_aiParseBench (tables)93.9%, +7 pts over Opus 5; beats Fable, Gemini, Astra@jerryjliu0Roboflow vision/detection”By far the best vision model from Anthropic”; now among the models ahead of Google on the Playground leaderboard@skalskip92, @skalskip92

Eval details and caveats:

  • Vals run settings. RSI was run in native Claude Code at max effort, with 1M context, 128K max output tokens, and temperature 1 (@ValsAI).

  • ParseBench caveats. The model still struggles on charts, formatting, and layout. At 5.8¢/page, LlamaIndex calls it too expensive for production OCR. That verdict comes from a vendor with a competing product.

  • AI R&D vs coding. @eliebakouch reads the system card as “roughly similar on AI R&D but a beast on agentic coding.”

  • Saturation. @scaling01 asked whether CoBench is “cooked.” @synthwavedd joked about a new benchmark that launched already saturated.

  • Arena. Opus 5.5 is in Agent Arena and in Battle Mode for WebDev, Text, Vision, and Document. No scores yet (@arena).

Effort‑scaling anomaly. On an agentic coding chart, xhigh effort costs about 2.8x more than medium for a 3.2‑point lower score (@LearnOpenCV). @Yuchenj_UW called it the “most bizarre benchmark result” and advised sticking with medium.

@nrehiew_ offered an explanation:

  • Opus 5 showed the same pattern on FrontierCode.

  • FrontierCode penalizes unnecessary changes, and higher effort produces scope creep.

  • As a result, models “consistently perform worse at higher reasoning efforts.”

System card details

  • Multi‑agent scaling. The system card reports scaling up to 100 parallel agents in Section 8.12. @scaling01 called it the first lab report of its kind. @maksym_andr highlighted it as evidence on multi-agent scaling laws.

  • ProgramBench caveats. ProgramBench author @OfirPress flagged that Anthropic’s near‑100% solve rate comes from a 166/200 subset. That subset likely excludes the hardest programs, such as FFmpeg and the PHP compiler. He also flagged a metric mismatch (@OfirPress, @OfirPress):

    • Anthropic reports average test pass rate.

    • ProgramBench reports full task completion.

    • Partial solves often pass 60–70% of tests, which inflates the pass-rate metric.

  • Comparison with Mythos 5.1. Opus 5.5 outscores Mythos 5.1 on Anthropic’s ECI and beats it on every tested cyber eval (@scaling01, @scaling01).

  • Odd misalignment finding. @teortaxesTex quoted a passage: malicious output occurred “almost exclusively in cases where, prior to the malicious output, Claude made an improbable, innocuous mistake.” He asked whether Anthropic had “sleeper-agent[ed] themselves.”

  • “Trained from RSI.” He separately quoted a line about “the first model trained from RSI” and called it concerning (@teortaxesTex).

  • Biomedical imaging. @iScienceLuvr welcomed the reported biomedical image analysis capabilities.

  • Requests for more. @scaling01 asked for time horizons without chain-of-thought.

Safety posture and safeguard controversy

Official position:

  • Sam Bowman: Opus 5.5 is “sufficiently safer than its predecessors that releasing it, more likely than not, reduces risks related to misalignment,” especially for the most extreme alignment risks (@sleepinyourhat, @sleepinyourhat).

  • He also acknowledged worry about keeping pace with escalating risk, while saying current tools remain trustworthy at this capability level (@sleepinyourhat).

  • Mike Krieger cited extensive alignment testing and outside evaluation, including by METR (@mikeyk).

Friction:

  • Over-triggering fallback. @iScienceLuvr got downgraded to the fallback model after asking Opus 5.5 to cure cancer.

  • China targeting (single test). @xlr8harder says a quick test suggests the frontier-LLM-development classifiers target Chinese hardware. He calls for more probing.

  • Reactions to the China angle. @teortaxesTex framed this as Anthropic undermining Chinese AI. @jakehalloran1 read it as protecting Trainium know‑how.

“Pacing the frontier” framing:

  • @theo argued none of today’s releases were Astra‑ or Fable‑tier and that this is deliberate pacing.

  • @goodside said lab calls to pace the frontier have weakened his “pause and do what?” stance.

  • @dejavucoder mocked the framing, given that Opus 5.5 outperforms Fable 5.1.

Writing, prompting, and behavior

  • Writing fixes from staff. “We fixed the writing” (@_sholtodouglas) and “we fixed the accent” (@NotTomBrown).

  • Unusual candor. @nmca (Anthropic) posted: “way, way, way better than Opus 5. Sorry about that model.” @theo called it a wild tweet that signals looser comms.

  • Em dashes. @theo reports they are gone from output. It was the most‑engaged reaction post.

  • Anthropic’s prompting playbook (@ClaudeDevs):

    • Hand over a whole task and define “done” and check‑in points.

    • Drop “think carefully,” since the model always thinks first.

    • After a long run, ask what it needs to go further.

  • Why old tricks break. @dbreunig notes old prompt tricks now clash with the model’s training, an argument for re‑compilable prompt optimization.

  • Long-run steering. @omarsar0 highlights Anthropic’s prompt for long runs, where the model sometimes stops to report instead of continuing.

  • Bug report. The live model sometimes generates user turns (@BlackHC).

  • Writing quality in practice. Hamel Husain livestreamed “Is Slop Dead?” testing its writing (@HamelHusain). @nptacek shared a one‑shot result from a personal writing eval.

Vision, 3D, and code-as-art demos

  • Improved perception. Sholto Douglas says the 5.5 series has “a serious step up” in 3D understanding and modeling, and that the model “can see now; it was a bit blind before” (@_sholtodouglas, @_sholtodouglas).

  • Painting in code. @jkeatn had the model generate paintings with pure Python, pixel by pixel:

    • About 7,500 lines of code using standard libraries to emulate brush styles.

    • No image model and no reference images.

    • Sholto contrasts this “manual brush” creativity with diffusion models (@_sholtodouglas).

  • Blender scenes. Alex Albert showed Blender claymations from one prompt on claude.ai (@alexalbert__). He also built a source‑grounded 1906 San Francisco Market Street:

    • Built from Sanborn maps, period film, and archival photos.

    • Procedural generators only, with no downloaded meshes or textures (@alexalbert__, prompt).

    • @karpathy riffed on the idea: turn historical images or video into custom GTA‑style worlds you can walk through.

  • More demos:

    • A code‑drawn JS animation (@kevin_t_ngo) and an official exploration thread (@claudeai).

    • A code‑generated Golden Gate Bridge, judged “as good as Astra” at 3D scenes (@petergyang).

    • “Best visual design of any model I’ve tested” (@other__reality).

    • A coral reef wallpaper; the builder says it feels about 3x faster and cheaper (@chaseleantj).

  • Open question. @teortaxesTex asks why this generation is so good at mapping functions to pixels, and suggests generalization.

Reactions: supportive, skeptical, comparative

Supportive:

  • Pipeline bugs. @rishdotblog says it found pipeline issues that Fable and Astra missed. It also found 7 SEC filing errors, including a Comfort Systems XBRL mis‑tag of Q1 revenue as full‑year (@rishdotblog).

  • Returning users. “Claude is back”: @Yuchenj_UW says he is returning to Claude Code after a month away.

  • Usage limits. Heavy all‑day use “barely making a dent” in limits (@theo).

  • Nostalgia. Comparisons to the well‑liked Opus 4.5 and 4.6 (@arohan, @kimmonismus).

  • Competitive framing. @scaling01 said Anthropic is “frontier‑mogging again.” @kimmonismus said “they chose war with OpenAI.”

Skeptical or neutral:

  • Trust deficit. @kylebrussell says he no longer trusts Opus releases to feel better. Sholto replied asking whether this one resets that trust (@_sholtodouglas).

  • Limits don’t matter to everyone. @stablequan never hits the limits anyway.

  • Price as headline. @dbreunig asked what it means that both labs’ headline feature is cheaper tokens.

Head-to-head with GPT‑6 Sol:

  • For Opus. @andrew_n_carr says Opus 5.5 “runs circles around” Sol. @synthwavedd says Sol came in below expectations and Anthropic “wins the day.”

  • Against. @teortaxesTex argues Opus 5.5’s cost and multi‑agent wins are “effectively negated with Astra+Sol+Luna spam,” since Sol is half Opus 5.5’s price (@scaling01).

  • Neutral. @kimmonismus‘s recap calls it no clear winner: Anthropic led on capability surprise, OpenAI on price. @simonw published a writeup comparing all three models with pelican grids across effort levels.

OpenAI’s GPT-6 Sol and Luna: Cheaper Astra-Derived Models for Codex, Work, and API

  • OpenAI answered within hours with GPT-6 Sol and GPT-6 Luna, described as faster, cheaper models that inherit much of GPT-6 Astra’s advances in coding, computer use, factuality, and alignment @OpenAI @OpenAIDevs. Pricing is aggressive: Sol at $2 / $10 per million input/output tokens and Luna at $0.10 / $0.50, each about 50% cheaper than their GPT-5.6 predecessors @OpenAI. They rolled out to ChatGPT Work and Codex plus the API, with Luna also available to Free and Go users in the desktop app, though notably not yet in Chat mode @OpenAI.

  • OpenAI’s comparison framing focused on cost-per-task Pareto gains rather than absolute flagship frontier wins. Their published examples claim Sol at xhigh effort beats Claude Opus 5 max on AutomationBench at roughly 9% of the cost per task, while Luna max exceeds GPT-5.6 Sol medium on OSWorld 2.0 offline at one-tenth the cost @reach_vb. Third-party integrations moved quickly: Perplexity made Sol its default “Light” effort orchestrator @perplexity_ai, Devin reported Sol matching GPT-5.6 Sol at 61% lower cost per task and Luna beating its predecessor at roughly a quarter of the cost @cognition, and Arena added both for agentic and code-side testing @arena.

  • The deeper infrastructure story may matter more than the SKU names. OpenAI said it improved caching and inference efficiency, exposing up to 90% discounts on cached input-token reads and a new Prompt Caching Dashboard plus diagnostics API to understand broken cache reuse @OpenAIDevs @OpenAIDevs. That’s particularly relevant for long-running agents where cache invalidation from tool toggles or reasoning changes has been costly. Market reaction was mixed: many praised the economics, especially Luna’s price floor, while others felt Anthropic won on headline model quality and OpenAI won on affordability and deployment ergonomics @kimmonismus @synthwavedd.

Agent Infrastructure, Eval Tooling, and Post-Training from Real Use

  • Several posts converged on a now-familiar pattern: value is shifting from raw model access to harnesses, evals, routing, and post-training on proprietary trajectories. DigitalOcean Managed Agents entered public preview with support for Claude Code, Codex, and LangGraph-style agents, plus pause-when-idle runtimes, governed tool endpoints, and 75+ model choices @digitalocean. On the developer workflow side, VS Code Agent Merge introduced an experimental mode for resolving review comments, failed checks, and merge conflicts automatically inside PRs @code.

  • Perplexity shared one of the more concrete post-training reports: its Computer agent uses a mix of rejection-sampling fine-tuning and hint-guided self-distillation on real user sessions to learn from successful trajectories and explicit tool-call mistakes, with a claimed 21.2% reduction in tool-call failures in a live A/B test @perplexity_ai @AravSrinivas. That’s a useful example of labs operationalizing sim-to-real bridging via production traces rather than purely synthetic RL environments.

  • Eval and observability tooling also got attention. Lenny’s newsletter highlighted concrete ROI from eval investment across companies like Ramp, Shopify, Harvey, and Cursor, and linked a sequel from Hamel Husain and Shreya Shankar on advanced eval systems @lennysan. Hamel also released an evals skill/plugin intended to automate parts of eval auditing and error analysis @lennysan. On the observability side, LangSmith shipped improved support for decision models like Jev/SemIf, making state, questions, choices, and outputs easier to inspect in agent traces @hwchase17 @LangChain. The meta-point from multiple practitioners: the harness can materially change benchmark outcomes even for the same model and prompt @omarsar0.

Open Models, Compression, and Systems Work for Running Bigger Models on Smaller Hardware

  • Tim Dettmers kicked off an “open-source week” with a runtime dynamic compression framework integrated into bitsandbytes2, targeting 1.5–2.0 bit compression at high quality and promising “lazy compression” that automatically finds a better memory/quality/speed tradeoff at deployment time @Tim_Dettmers @Tim_Dettmers. The framing is explicitly for small teams and individuals trying to run large open-weight models with constrained memory, including growing KV caches.

  • Hardware and local deployment were another theme. A hands-on post about NVIDIA DGX Spark described a 15×15×5.05 cm, 1.2 kg system with GB10 Grace Blackwell, 128 GB unified memory, and up to 1 PFLOP FP4 sparse theoretical compute, with NVIDIA claiming support for local inference on models up to 200B parameters with quantization and fine-tuning up to 70B with methods like QLoRA @kimmonismus. Separately, Reka EdgeQ showed an on-device VLM optimized directly for Qualcomm’s Hexagon NPU, with 0.73s TTFT, 6.9 mWh per inference, and the GPU kept idle for sustained thermal performance @RekaAILabs.

  • On the open-model side, there were a few meaningful releases rather than just commentary. Step Code v0.1.0 launched under MIT, packaging a coding agent CLI with reported scores of 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark @StepFun_ai. Ming-Image-0.1-Design, a 6B open-weight image design family, was released alongside UI-design and image-to-editable-PPT “agent skills,” claiming #1 among open-weight models on Artificial Analysis’s UI/UX design leaderboard @AntLingAGI. A smaller but technically notable pretraining result came from Rigel, a 2.3B MoE / 360M active Hybrid Mamba-2 reportedly trained across mixed H100/A100/V100 and TPU v5p/v6e hardware on one codebase, reaching within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs @MayankMish98.

Multimodal Models: Image, Video, Speech, and World Models

  • In image generation and editing, Qwen-Image-2.1 had a strong day on community leaderboards, taking #1 among open models in both the Image Edit Arena and Text-to-Image Arena, landing close to frontier proprietary systems overall @arena. Supporting ecosystem work included Unsloth Desktop support with INT8/FP8 and GGUFs that can fit under 6–8 GB VRAM with RAM offloading @danielhanchen, and Gradio’s effort to shrink Qwen’s default 9B prompt rewriter down to 0.8B for laptop use @Gradio.

  • In video and real-time media, PixVerse R2 was announced as a real-time world model emphasizing editable, persistent “living worlds” @PixVerse, while fal published a stack breakdown for H3 Max, claiming 5 seconds of video generated in 3 seconds through optimizations spanning post-training, GPU execution, weight loading, scaling, and serving @fal. Their World Model Accelerator interface is notable for replacing request/response semantics with a persistent WebRTC session for interactive models @fal.

  • Speech remained active too. AssemblyAI Universal-3.5 Pro went live on OpenRouter with 19-language synchronous STT, domain steering via keyterms and prompting, and a temporary discount @OpenRouter. StepAudio 3 ASR reached 1.7% WER on the Artificial Analysis AA-WER Index, essentially tying the top spot for non-streaming speech-to-text, albeit at a premium price @ArtificialAnlys. Moondream also released Parakeet Redux and Parakeet Ultra local STT models for 25 languages, targeting CPU and GPU respectively @moondreamai.

Top tweets (by engagement)

  • Claude Opus 5.5 launch: Anthropic’s main release post dominated engagement and framed the day’s biggest model event @claudeai.

  • GPT-6 Sol and Luna launch: OpenAI’s release of cheaper Astra-derived models was the other major headline @OpenAI.

  • Managed Agents preview: DigitalOcean’s public preview of managed runtimes for Claude Code/Codex/custom agents drew unusually high infra interest @digitalocean.

  • OpenAI standards proposal: Sam Altman’s post on AI standards and governance generated heavy discussion beyond pure product news @sama.

  • Epoch on AI cost curves: Epoch’s estimate that AI cost at fixed performance has been falling ~47% per quarter since 2023 was one of the more useful macro datapoints of the day @EpochAIResearch.


AI Reddit Recap

/r/LocalLlama + /r/localLLM Recap

1. Qwen, DeepSeek and AliceAI Large-Model Roadmaps

  • Qwen4-27B just confirmed (Activity: 2353): A conference slide image appears to confirm an upcoming Qwen4 Series lineup, explicitly listing Qwen4-27B alongside Qwen4-Max, Qwen4-Flash, and Qwen4-Plus. The post highlights community interest in whether Alibaba will also release a smaller MoE-style variant like 35B-A3B, and commenters speculate that architectural changes such as N-grams could reduce VRAM requirements. Commenters are mainly debating whether Qwen4-27B will outperform Qwen 3.8 Flash Next and whether the best local inference path will favor discrete GPUs or high-capacity unified-memory systems. There is also interest in comparing Qwen4 Flash, Qwen3.8 Flash Next, and Qwen4-27B if all are released as open weights.

    • Commenters speculated that Qwen4-27B could have lower VRAM requirements if it adopts an N-gram-style architecture, though no concrete implementation details or memory figures were provided in the thread.

    • A technical comparison was proposed between Qwen4 Flash, Qwen3.8 Flash Next, and Qwen4-27B, assuming all are released as open weights. The key question raised was whether a dense/standard 27B model would outperform a smaller Flash variant enough to influence whether users prioritize discrete GPUs or large unified-memory systems.

    • One user hoped Qwen4 Flash retains the memory footprint of Flash Next, specifically targeting deployment within 128 GB of VRAM, implying interest in local inference feasibility for larger open-weight Qwen models.

  • Alibaba plans AI model with 5 trillion to 10 trillion parameters, unveils new chip (Activity: 648): Alibaba reportedly plans an AI model in the 5T–10T parameter range and unveiled a new AI chip, implying a frontier-scale training/inference target far beyond current consumer/local deployment practicality. Commenters contextualize this against prior excitement around DeepSeek R1’s 671B/691B-class scale and expect any practical downstream use to come via distillation into smaller Qwen-family models such as a hypothetical Qwen 4 27B. The main debate is skepticism about local inference feasibility—“minutes per token”—versus optimism that Alibaba may distill a much larger internal model, possibly “Astra,” into a genuinely competitive Chinese frontier model.

    • Commenters noted that a 5T–10T parameter Alibaba model would be effectively API-only for almost all users, with local inference on homelab hardware being impractical and potentially yielding extremely slow minutes-per-token generation without major sparsity, quantization, or specialized serving hardware.

    • Several comments framed the likely practical value as distillation, comparing it to the excitement around DeepSeek R1’s 671B/691B-class parameter count and suggesting users may instead wait for a smaller descendant such as a hypothetical Qwen 4 27B that could run locally.

    • One commenter speculated that if Alibaba has successfully distilled or incorporated capabilities from Astra, it could indicate a more serious Chinese frontier-model push, though the thread provides no benchmark evidence or implementation details to validate that claim.

Keep reading with a 7-day free trial

Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Latent.Space · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture