<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Latent.Space]]></title><description><![CDATA[The AI Engineer newsletter + Top technical AI podcast. How leading labs build Agents, Models, Infra, & AI for Science. See https://latent.space/about for highlights from Greg Brockman, Andrej Karpathy, George Hotz, Simon Willison, Soumith Chintala et al!]]></description><link>https://www.latent.space</link><image><url>https://substackcdn.com/image/fetch/$s_!DbYa!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b0838a-bd14-46a1-801c-b6a2046e5c1e_1130x1130.png</url><title>Latent.Space</title><link>https://www.latent.space</link></image><generator>Substack</generator><lastBuildDate>Tue, 14 Jul 2026 13:24:19 GMT</lastBuildDate><atom:link href="https://www.latent.space/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Latent.Space]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[swyx@noreply.com]]></webMaster><itunes:owner><itunes:email><![CDATA[swyx@noreply.com]]></itunes:email><itunes:name><![CDATA[Latent.Space]]></itunes:name></itunes:owner><itunes:author><![CDATA[Latent.Space]]></itunes:author><googleplay:owner><![CDATA[swyx@noreply.com]]></googleplay:owner><googleplay:email><![CDATA[swyx@noreply.com]]></googleplay:email><googleplay:author><![CDATA[Latent.Space]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[[AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??]]></title><description><![CDATA[a quiet day lets us fact check some numbers against the sound of silence of Claude Code reporting...]]></description><link>https://www.latent.space/p/ainews-codex-usage-up-10x-in-6-months</link><guid isPermaLink="false">https://www.latent.space/p/ainews-codex-usage-up-10x-in-6-months</guid><pubDate>Tue, 14 Jul 2026 01:22:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!cqvt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Congrats to Allen for the <a href="https://www.youtube.com/watch?v=jhpmMTus5a0">next episode of the Latent Space Food show with Engram CEO Dan Biderman today</a>, and to <a href="https://www.youtube.com/watch?v=V-EDrhIhHzQ&amp;t=1s">the Prime Intellect folks on their 1B valuation, $100M ARR, and verifiers v1</a>.</p><p>Today was pretty quiet and people are still deeply digesting <a href="https://www.latent.space/p/ainews-not-much-happened-today-f5c">last week&#8217;s multiple frontier model launches</a>. We were going to write &#8220;not much happened today&#8221;, but we also have <a href="https://www.latent.space/p/ainews-sci-fi-with-a-touch-of-madness?utm_source=publication-search">a policy of updating you repeatedly on outlier trends</a> that you should really be on top of. In reviewing the Reddit AINews recaps below surfaced <a href="https://www.reddit.com/r/ClaudeCode/comments/1uuqz4l/anthropic_i_think_you_really_need_to_react_youre/">this post</a>, we saw a tweet we had missed before - </p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/thsottiaux/status/2076365965915467978&quot;,&quot;full_text&quot;:&quot;Morning. The last 48 hours of Codex and ChatGPT Work have been intense! Three important updates:\n\n- Temporarily removing the 5 hour usage limit restriction for all Plus, Business and Pro plans\n- Rolling out changes that will make GPT 5.6 Sol more efficient across the board and&quot;,&quot;username&quot;:&quot;thsottiaux&quot;,&quot;name&quot;:&quot;Tibo&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2075819673263001600/pj1vyX6I_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-12T17:59:57.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:2796,&quot;retweet_count&quot;:1947,&quot;like_count&quot;:25058,&quot;impression_count&quot;:4294105,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p><a href="https://www.latent.space/p/ainews-openai-launches-gpt-56-solterraluna">GPT 5.6 was launched on July 9</a>. </p><p>This tweet on July 12 says they hit 6M users in the prior 48 hours (Jul 10-12).</p><p>Then 24.5 hours later Tibo reports 7M users&#8230;</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/thsottiaux/status/2076735790567338203&quot;,&quot;full_text&quot;:&quot;Thank you to the 7M active users who are now using Codex and ChatGPT Work.\n\nWe have added a banked reset to everyone's account to celebrate the milestone. You can apply the reset in the desktop app or on web and it will replenish the weekly usage for you.\n\nHave fun out there.&quot;,&quot;username&quot;:&quot;thsottiaux&quot;,&quot;name&quot;:&quot;Tibo&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2075819673263001600/pj1vyX6I_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-13T18:29:31.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:1377,&quot;retweet_count&quot;:656,&quot;like_count&quot;:14101,&quot;impression_count&quot;:946943,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>&#8230;oddly coinciding with <a href="https://x.com/claudeai/status/2076351399999557669?s=20">a surprise extension of Claude Fable&#8217;s subscription status</a> (we have of course no idea if the two are related, but the permanently online conspiracy theorists are of course making a connection).</p><p>We of course recall Fidji&#8217;s <a href="https://x.com/fidjissimo/status/2033537381907710092">March disclosure of 2M Codex users</a>, which allows us to update our <a href="https://www.youtube.com/watch?v=5N33E9tC400&amp;t=401s">AIE NYC 2025</a> chart (<a href="http://ai.engineer/nyc">AIE NYC 2026</a> is next!):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cqvt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cqvt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 424w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 848w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 1272w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cqvt!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png" width="1200" height="779.8270893371758" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:902,&quot;width&quot;:1388,&quot;resizeWidth&quot;:1200,&quot;bytes&quot;:235691,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206943176?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cqvt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 424w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 848w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 1272w, https://substackcdn.com/image/fetch/$s_!cqvt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Comparatively, the last update we got about Claude Code is the <a href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation">roughly 2M users and $2.5B ARR in Feb</a> (&#8220;The number of weekly active Claude Code users has also doubled since January 1 [six weeks ago]."). Now we have a sense of where Codex started the year (Fidji <a href="https://x.com/fidjissimo/status/2033537381907710092">puts the Jan 1 number at around 550k-700k users</a>), we can reasonably conclude that Codex has followed a similar trajectory and is now around 10x user growth year to date.</p><p>The charitable interpretation on Claude Code&#8217;s comparative silence on reporting, of course, is that <a href="https://www.latent.space/p/ainews-claude-tag-multiplayer-proactive?utm_source=publication-search">they moved the bulk of coding to Claude Tag months ago and are now focusing users there</a>, which will have different/hard to compare usage statistics given the different accessibility of a Slackbot vs a CLI tool. </p><p>But 10x growth in 6 months is an impressive number to beat nonetheless.</p><p></p><blockquote><p>AI News for 7/11/2026-7/13/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Agent RL Infrastructure: Prime Intellect&#8217;s Verifiers v1 and Long-Horizon Rollouts</strong></p><ul><li><p><strong>Prime Intellect&#8217;s verifiers v1</strong>: <a href="https://x.com/PrimeIntellect/status/2076447247693402301">Prime Intellect</a> released <strong>verifiers v1</strong>, a substantial redesign of its environment stack for <strong>agentic RL and evals</strong>. The key abstraction splits environments into a <strong>taskset, harness, and runtime</strong>, explicitly supporting &#8220;bring your own harness&#8221; workflows for coding and computer-use agents across heterogeneous execution setups, as highlighted by <a href="https://x.com/johannes_hage/status/2076447852528889939">Johannes Hage</a> and in a <a href="https://x.com/johannes_hage/status/2076449075621462457">follow-up deep dive</a>. The release was framed by team members as months of infra modernization work with major efficiency gains, including richer commentary from <a href="https://x.com/willccbb/status/2076449433483616346">willccbb</a>, <a href="https://x.com/mikasenghaas/status/2076507323561021779">mikasenghaas</a>, and <a href="https://x.com/xeophon/status/2076509926256422947">xeophon</a>.</p></li><li><p><strong>Why it matters technically</strong>: one of the most important underlying changes is that rollout traces are now stored as <strong>message DAGs</strong>, so each message is stored once instead of repeatedly copied into full histories; that shifts trace growth from <strong>O(n&#178;)</strong> to <strong>O(n)</strong> in turn count, making long-horizon multimodal rollouts and router replay much more practical, per <a href="https://x.com/PrimeIntellect/status/2076447253938786648">Prime Intellect</a>. The team also claimed a concrete training configuration: a <strong>100B reasoning model</strong>, on <strong>40-turn SWE agent tasks</strong>, in a user-supplied coding harness, for <strong>1000 RL steps</strong>, using <strong>6 H200 nodes</strong> in <strong>under 2 days</strong> (<a href="https://x.com/willccbb/status/2076451043504967783">willccbb</a>). That claim was reinforced by ecosystem support from <a href="https://x.com/vllm_project/status/2076528386927997249">vLLM</a>, which noted verifiers&#8217; rollout path runs on vLLM with exact token IDs/logprobs to avoid tokenization drift between serving and training.</p></li></ul><p><strong>Coding Agents, Harness Design, and Cost-Per-Task Competition</strong></p><ul><li><p><strong>Harnesses are becoming the product surface</strong>: several posts converged on the idea that model quality is no longer the only differentiator; the <strong>harness/orchestrator</strong> increasingly determines outcomes. <a href="https://x.com/localfirstconf/status/2076678392615682215">threepointone&#8217;s talk</a> was summarized as &#8220;the harness is the app,&#8221; while <a href="https://x.com/hwchase17/status/2076784403414651035">LangChain</a> argued that winning agent products will come from <strong>task-specialized harnesses</strong>, not generic wrappers. <a href="https://x.com/FactoryAI/status/2076710400729731349">Factory</a> pushed a related UI angle with &#8220;design mode,&#8221; where users point at UI elements/files instead of verbally re-specifying edits. On the orchestration side, <a href="https://x.com/omarsar0/status/2076720090549035318">omarsar0</a> emphasized provider-switching across models as a hedge against pricing/policy churn.</p></li><li><p><strong>Benchmarks are moving from token price to cost per task</strong>: <a href="https://x.com/skirano/status/2076456519810580681">skirano</a> built a coding-agent index explorer and found notable cost/perf tradeoffs such as <strong>Terra Max slightly ahead of Fable 5 Max</strong> on score for materially lower cost, while <a href="https://x.com/cognition/status/2076714965344342382">Cognition</a> reported that <strong>Devin Fusion</strong> now uses <strong>Fable 5</strong> and that, surprisingly, it can be <strong>lower cost per task than Opus 4.8</strong> because stronger delegation and judgment reduce unnecessary work. <a href="https://x.com/imjaredz/status/2076715750715482162">imjaredz</a> highlighted the key stat from those experiments: in <strong>81% of Fable-led runs</strong>, the lead model never makes a code edit, implying expensive models can be cheaper when they avoid wasted actions.</p></li><li><p><strong>Real-world agent benchmarks are getting denser</strong>: <a href="https://x.com/arena/status/2076709326711037991">Arena</a> placed <strong>GPT-5.6 Sol</strong> at <strong>#2</strong> on its agent leaderboard based on <strong>7.8K real-world agentic sessions</strong>, with strong steerability and task success; later, <a href="https://x.com/arena/status/2076728509813469536">Arena</a> put <strong>Grok-4.5</strong> at <strong>#13</strong>, a significant jump over Grok 4.3. <a href="https://x.com/ArtificialAnlys/status/2076791491071295708">Artificial Analysis</a> also emphasized <strong>cost per task</strong> as an increasingly important metric for long-horizon knowledge work, arguing token pricing alone misses effects from turns, verbosity, and cache hit rates. Separate evaluation work from <a href="https://x.com/doesdatmaksense/status/2076642415767965701">Parlance Labs</a> compared automated eval platforms and foundation models on failure analysis over production voice-agent traces, while <a href="https://x.com/dair_ai/status/2076699431207154069">dair.ai</a> highlighted a paper on the <strong>anatomy of CLI coding-agent failures</strong>, focusing on where runs become unrecoverable rather than only final pass/fail.</p></li></ul><p><strong>OpenAI GPT-5.6 Sol, Codex Usage Fixes, and Product Surface Expansion</strong></p><ul><li><p><strong>OpenAI addressed Codex/Sol usage burn transparently</strong>: the biggest operational thread came from <a href="https://x.com/thsottiaux/status/2076495156757577895">thsottiaux</a>, who explained several fixes for <strong>GPT-5.6 Sol</strong> in ChatGPT Work/Codex: inference optimizations yielding roughly <strong>10% more usage</strong>, a rollback of context limit from <strong>372k</strong> to <strong>272k</strong> after billing/usage side effects, reversion of some experimental reasoning-effort (&#8220;<strong>juice</strong>&#8221;) changes, and fixes for overactive multi-agent behavior at high/xhigh settings. Community reverse-engineering from <a href="https://x.com/theo/status/2076512403668488299">theo</a> proposed that compounding factors around long context, subagent spawning, and fast mode were behind the severe burn, though he later corrected one billing detail in a <a href="https://x.com/theo/status/2076543971216830551">follow-up</a>. Reactions split between criticism of a perceived &#8220;nerf&#8221; narrative (<a href="https://x.com/ns123abc/status/2076498300312703349">ns123abc</a>) and praise for unusual transparency (<a href="https://x.com/theo/status/2076501402822775267">theo</a>, <a href="https://x.com/sama/status/2076696938918084809">sama</a>).</p></li><li><p><strong>Users are reporting strong coding/computer-use capability</strong>: multiple practitioners argued that <strong>OpenAI has taken the lead on coding models</strong>, including <a href="https://x.com/schrockn/status/2076488446961709218">schrockn</a>, while <a href="https://x.com/gdb/status/2076518764112445861">gdb</a> repeatedly showcased <strong>ChatGPT Work</strong> and Codex workflows for startup prospecting, web design, mobile work, and site generation. Particularly illustrative user demos included <a href="https://x.com/Star_Knight12/status/2076631428926972177">Star_Knight12</a> using <strong>Sol in Cursor</strong> to set up Blender MCP and render a floating MacBook without prior Blender experience, and <a href="https://x.com/petergostev/status/2076692164310884468">petergostev</a> showing <strong>GPT-5.6 Sol Ultra</strong> building a <strong>Doom-like game in SQL</strong>.</p></li><li><p><strong>Product-level expansion continues</strong>: <a href="https://x.com/ChatGPTapp/status/2076654365121855835">ChatGPTapp</a> announced ChatGPT&#8217;s return to <strong>WhatsApp in the EEA</strong>, plus Kakao/Viber support in additional markets. <a href="https://x.com/OpenAIDevs/status/2076715478878474575">OpenAIDevs</a> opened submissions for <strong>OpenAI Build Week</strong>. Across the OpenAI ecosystem, <a href="https://x.com/gdb/status/2076685930002538875">gdb</a> summarized the moment succinctly: &#8220;you can just create things.&#8221;</p></li></ul><p><strong>Open Models, Inference Systems, and Quantization</strong></p><ul><li><p><strong>Transformers&#8596;vLLM integration removes duplicated model implementation work</strong>: <a href="https://x.com/ClementDelangue/status/2076763231788339669">Clement Delangue</a> highlighted a major open-inference usability improvement: <strong>Hugging Face Transformers models can now run in vLLM at native speed</strong>, often matching or exceeding hand-written implementations. If this generalizes broadly, it reduces the long-standing burden of implementing each new architecture twice&#8212;once for research/training and once for high-performance serving&#8212;and could materially accelerate adoption of new open model architectures.</p></li><li><p><strong>Quantization remains a major lever</strong>: <a href="https://x.com/waterloo_intern/status/2076460984475263401">waterloo_intern</a> previewed a new quantization method claimed to beat existing approaches, including NVIDIA&#8217;s ModelOpt, by finding better layerwise precision assignments <strong>faster</strong>, with <strong>more aggressive quantization</strong> and <strong>higher benchmark scores</strong>. Complementing that, <a href="https://x.com/UnslothAI/status/2076665500294394109">Unsloth</a> published an AWS guide to <strong>LLM quantization and deployment</strong> spanning GGUF, NVFP4, and FP8. There was also practitioner commentary around <strong>fp4 RL / fp4 serving</strong> from <a href="https://x.com/nrehiew_/status/2076654135559233857">nrehiew_</a>, arguing low-bit post-training may enable cheap serving with limited quality loss.</p></li><li><p><strong>GLM-5.2 and local/open coding stacks continue to gain traction</strong>: several users described moving real workflows onto open or semi-open setups. <a href="https://x.com/juanjucm/status/2076714987569963508">juanjucm</a> wrote up using <strong>GLM-5.2</strong> for coding-agent workflows, while <a href="https://x.com/TheZachMueller/status/2076746035758502275">TheZachMueller</a> reported migrating one actual work pipeline from Claude to a stack built around <strong>GLM 5.2 NVFP4</strong> plus <strong>Kimi K2.7 Code NVFP4</strong> on an <strong>8xB200</strong> node, getting denser reports for pennies albeit at slower wall-clock latency. <a href="https://x.com/nutlope/status/2076722464671793184">nutlope</a> also released <strong>LlamaCoder v4</strong>, rebuilt around GLM 5.2.</p></li></ul><p><strong>Security, Privacy, and Data Control in Agent Tooling</strong></p><ul><li><p><strong>Grok Build code upload controversy</strong>: the most consequential security story came from <a href="https://x.com/IntCyberDigest/status/2076689215258014069">IntCyberDigest</a> and <a href="https://x.com/hrkrshnn/status/2076716354754015368">hrkrshnn</a>, who alleged that <strong>xAI&#8217;s Grok Build CLI</strong> was uploading entire repositories&#8212;including private code and secrets&#8212;to a Google Cloud bucket, far beyond what was needed for the coding task. The criticism centered on scope, silent server-side mitigation, and unclear retention/deletion guarantees. This triggered broader discussion about what agent tools actually transmit and why opt-out UX can diverge from wire-level behavior.</p></li><li><p><strong>xAI&#8217;s response emphasized ZDR and privacy controls</strong>: <a href="https://x.com/SpaceXAI/status/2076692402442846289#m">SpaceXAI</a> replied that for teams using <strong>zero data retention</strong>, trace and code data is not retained, API key use respects ZDR, and the <code>/privacy</code> command can disable retention and delete previously synced data. That answered some operational questions but did not fully resolve community concern around default behavior, prior uploads, and disclosure norms.</p></li><li><p><strong>Trust boundaries are becoming a central open-vs-closed argument</strong>: several posts extended the conversation beyond this incident. <a href="https://x.com/mchiang0610/status/2076736707471556755">mchiang0610</a> and <a href="https://x.com/jmorgan/status/2076750580052369896">jmorgan</a> argued that open models are not just about cost but about <strong>control over the human-AI learning loop</strong> and keeping institutional knowledge in-house. <a href="https://x.com/AravSrinivas/status/2076699450177892354">Arav Srinivas</a> said <strong>ZDR availability</strong> was one reason Perplexity integrated <strong>Grok 4.5</strong> quickly into its Computer harness.</p></li></ul><p><strong>Continual Learning, Multimodal Systems, and Research Directions</strong></p><ul><li><p><strong>Continual learning is re-emerging as a first-class systems problem</strong>: <a href="https://x.com/ysu_nlp/status/2076481232117067894">ysu_nlp</a> argued that a world where every organization owns its own human-AI learning loop depends on solving <strong>continual learning</strong>, and that current approaches&#8212;memory/RAG, domain post-training, task RL&#8212;are not yet sufficient. That theme recurred in new work from <a href="https://x.com/skyfallai/status/2076713589788864920">skyfallai</a>, which introduced <strong>Morpheus</strong>, described as a persistent enterprise simulation for real-world RL where the world does not reset; <a href="https://x.com/fchollet/status/2076719958189613307">fchollet</a> endorsed it as a benchmark better aligned with real deployment than stationary episodic RL.</p></li><li><p><strong>&#8220;Sleep and dreaming&#8221; for LLMs</strong>: <a href="https://x.com/behrouz_ali/status/2076710744456892519">behrouz_ali</a> and coauthors proposed that LLMs may need a <strong>sleep phase</strong> to consolidate short-term into long-term memory plus a <strong>dreaming phase</strong> for recursive self-improvement, introducing <strong>Knowledge Seeding</strong> and reporting benefits on continual learning/reasoning tasks. This dovetails with broader dissatisfaction around current continual-learning recipes and with <a href="https://x.com/kjaved_/status/2076663868160459214">Oak Lab</a>, the new venture from Rich Sutton and collaborators pursuing <strong>animal-like intelligence</strong> that learns from experience rather than today&#8217;s standard LLM pipeline.</p></li><li><p><strong>A broad spread of non-LLM-agent research shipped</strong>: notable items included <a href="https://x.com/SakanaAILabs/status/2076597965804765283">Sakana AI&#8217;s Smart Cellular Bricks</a> for decentralized physical self-recognition and repair in modular systems; <a href="https://x.com/HuggingPapers/status/2076513044340097501">ByteDance&#8217;s UniVR-34B</a>, described as learning reasoning/dynamics/planning directly from visual demonstrations; <a href="https://x.com/GoogleDeepMind/status/2076686114631340046">Google DeepMind&#8217;s Predicting the Past skill</a> for historical inference workflows; and <a href="https://x.com/AnthropicAI/status/2076719540785012872">Anthropic&#8217;s research</a> on how <strong>Claude&#8217;s expressed values</strong> vary across models and languages based on analysis of <strong>300K+ anonymized conversations</strong>.</p></li></ul><p><strong>Top tweets (by engagement)</strong></p><ul><li><p><strong>OpenAI Codex/Sol usage fixes</strong>: <a href="https://x.com/thsottiaux/status/2076495156757577895">thsottiaux on GPT-5.6 Sol usage, context, &#8220;juice,&#8221; and multi-agent fixes</a></p></li><li><p><strong>Grok Build privacy incident</strong>: <a href="https://x.com/IntCyberDigest/status/2076689215258014069">IntCyberDigest on full-repo uploads to xAI cloud buckets</a></p></li><li><p><strong>OpenAI response tone and user treatment</strong>: <a href="https://x.com/sama/status/2076780425280954658">sama: &#8220;come for the best model, stay because we don&#8217;t treat you with contempt&#8221;</a></p></li><li><p><strong>Prime Intellect rollout efficiency</strong>: <a href="https://x.com/willccbb/status/2076451043504967783">willccbb on training a 100B reasoning model for 40-turn SWE RL on 6 H200s in under 2 days</a></p></li><li><p><strong>Anthropic values research</strong>: <a href="https://x.com/AnthropicAI/status/2076719540785012872">Anthropic on model/language-dependent value expression across 300K+ conversations</a></p></li><li><p><strong>Transformers + vLLM interoperability</strong>: <a href="https://x.com/ClementDelangue/status/2076763231788339669">Clement Delangue on running Transformers models in vLLM at native speed</a></p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. E-Waste GPU Inference Benchmarks and Fixes</strong></h3><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-codex-usage-up-10x-in-6-months">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[[AINews] not much happened today]]></title><description><![CDATA[a quiet day after a week of nonstop model releases]]></description><link>https://www.latent.space/p/ainews-not-much-happened-today-f5c</link><guid isPermaLink="false">https://www.latent.space/p/ainews-not-much-happened-today-f5c</guid><pubDate>Sat, 11 Jul 2026 02:53:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7odD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>So <a href="https://x.com/swyx/status/2072961609094562283">dancing bugs</a> got upstaged by <a href="https://x.com/JangLawrenceK/status/2075204015890325703">kpop girls</a>, there&#8217;s the whole <a href="https://x.com/kellabyte/status/2075455336408871176">Bun vs Zig drama</a>, and <a href="https://www.latent.space/p/ainews-openai-launches-gpt-56-solterraluna">yesterday&#8217;s ChatGPT/Codex superapp launch</a> was <a href="https://x.com/thsottiaux/status/2075641131002700120">bumpier than expected</a>, and the <a href="https://x.com/steipete/status/2072061089177539003?s=46">reset button</a> was pressed a couple times to compensate.</p><p>After <a href="https://www.statsig.com/blog/openai-acquisition">buying Statsig</a> and making a big deal out of GPT5&#8217;s routing/getting rid of <a href="https://x.com/michpokrass/status/1922733011042250770">the model picker</a>, the main issue now is that GPT 5.6&#8217;s extra options are confusing people a bit. Most people just have a single slider:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fdhH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fdhH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 424w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 848w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 1272w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fdhH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png" width="420" height="274" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/780723c6-1158-4a6b-a957-490705fdba08_420x274.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:274,&quot;width&quot;:420,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19306,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206529076?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fdhH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 424w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 848w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 1272w, https://substackcdn.com/image/fetch/$s_!fdhH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F780723c6-1158-4a6b-a957-490705fdba08_420x274.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>But API users have literally 36 variants of GPT 5.6 now:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7odD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7odD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 424w, https://substackcdn.com/image/fetch/$s_!7odD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 848w, https://substackcdn.com/image/fetch/$s_!7odD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 1272w, https://substackcdn.com/image/fetch/$s_!7odD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7odD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png" width="1328" height="982" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:982,&quot;width&quot;:1328,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:552325,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206529076?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7odD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 424w, https://substackcdn.com/image/fetch/$s_!7odD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 848w, https://substackcdn.com/image/fetch/$s_!7odD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 1272w, https://substackcdn.com/image/fetch/$s_!7odD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Most people can get by with just 3 rough clusters</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/jumperz/status/2075618148133556421&quot;,&quot;full_text&quot;:&quot;so this is what i've found works best with gpt-5.6 so far.. \n\n&amp;gt; luna high, normal everyday coding, fast, capable, doesn't feel wasteful...\n\n&amp;gt;luna xhigh  better quality without jumping to the expensive models...\n\n&amp;gt;terra medium, bigger features\n\n&amp;gt;terra high, repo-wide changes..\n\n&amp;gt;&quot;,&quot;username&quot;:&quot;jumperz&quot;,&quot;name&quot;:&quot;JUMPERZ&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2066609443194773504/sSlzoKn2_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-10T16:28:24.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HM4TsU1XYAAMV_R.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/CvBee4zCqq&quot;}],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;gpt 5.6 is honestly making the $200 pro plan harder to justify&#8230;\n\nwhen 5.5 never made me think about usage.. \n\nnow you&#8217;ve got sol, terra, and luna, all with different limits and usage costs&#8230;\n\nso Instead of just picking the best model for the job, you&#8217;re constantly trying to make https://t.co/eA50elhmLP&quot;,&quot;username&quot;:&quot;jumperz&quot;,&quot;name&quot;:&quot;JUMPERZ&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2066609443194773504/sSlzoKn2_normal.jpg&quot;},&quot;reply_count&quot;:20,&quot;retweet_count&quot;:18,&quot;like_count&quot;:284,&quot;impression_count&quot;:31533,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>And many guides are coming up:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/rasbt/status/2075573860796436626&quot;,&quot;full_text&quot;:&quot;For agentic coding, one can say:\n\n- Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper).\n\n- Forget everything below Sol High, use Luna with higher effort settings here\n\n- Forget Sol Extra &quot;,&quot;username&quot;:&quot;rasbt&quot;,&quot;name&quot;:&quot;Sebastian Raschka&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1661187442043486209/a3E4t1eV_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-10T13:32:25.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HM3raXqWkAAMqGb.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/Tjc9mELCer&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:250,&quot;retweet_count&quot;:272,&quot;like_count&quot;:3108,&quot;impression_count&quot;:458705,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p></p><p>The top AIE talk so far this week has been Theo&#8217;s closing keynote, and the last of the online track will be released this weekend.</p><div id="youtube2-xUnRQ9vLXxo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;xUnRQ9vLXxo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/xUnRQ9vLXxo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><blockquote><p>AI News for 7/09/2026-7/10/2026. We checked 12 subreddits and <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a>. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>OpenAI&#8217;s GPT-5.6 rollout: model stratification, agent UX, and early benchmark signals</strong></p><ul><li><p><strong>GPT-5.6 introduced a more explicit model/compute ladder</strong>: users are now navigating <strong>Luna / Terra / Sol</strong> plus multiple effort levels, with community guidance converging around &#8220;start lower than you did on 5.5.&#8221; OpenAI staff explained that <strong>Max</strong> means one model spending longer on a hard problem, while <strong>Ultra</strong> parallelizes work across subagents; they also noted that 5.5&#8594;5.6 effort settings are <strong>not directly comparable</strong> (<a href="https://x.com/reach_vb/status/2075489301253488778">guidance from @reach_vb</a>, <a href="https://x.com/pvncher/status/2075590107214520590">follow-up</a>, <a href="https://x.com/gabrielchua/status/2075521933576462357">practical default suggestion</a>). The community reaction was mixed: many praised the added control, while others criticized the <strong>30+ configuration combinatorics</strong> and missing &#8220;Auto&#8221; routing (<a href="https://x.com/rasbt/status/2075369179817902176">@rasbt</a>, <a href="https://x.com/Yuchenj_UW/status/2075627844412264796">@Yuchenj_UW</a>).</p></li><li><p><strong>The product launch landed with real UX regressions, and OpenAI publicly course-corrected fast</strong>: users complained that the new <strong>ChatGPT Work / Codex</strong> split was confusing, chats/projects became harder to find, and usage burned down faster than expected (<a href="https://x.com/scaling01/status/2075595915419599176">@scaling01</a>, <a href="https://x.com/simonw/status/2075663372323008755">@simonw</a>, <a href="https://x.com/kimmonismus/status/2075608495756333087">@kimmonismus</a>). OpenAI responded unusually directly: <strong>multiple usage-limit resets</strong>, acknowledgements that defaults nudged users toward overly expensive settings, and a commitment to restore familiar sidebar/navigation patterns and clarify positioning between Work and Codex (<a href="https://x.com/thsottiaux/status/2075452680760443190">@thsottiaux reset announcement</a>, <a href="https://x.com/reach_vb/status/2075460193681367532">second reset</a>, <a href="https://x.com/thsottiaux/status/2075641131002700120">full corrective roadmap</a>).</p></li><li><p><strong>Initial eval picture</strong>: GPT-5.6 appears strongest in <strong>agentic coding / presentation / some science tasks</strong>, but not unambiguously dominant everywhere. Examples: <strong>#1 tie in Code Arena: Frontend</strong> with Claude Fable 5 while being ~<strong>2&#215; cheaper</strong> on listed IO pricing (<a href="https://x.com/arena/status/2075672492312768683">Arena</a>); best recorded <strong>Presentation Elo</strong> on AA-Briefcase with a ~<strong>500-point</strong> jump over GPT-5.5 (<a href="https://x.com/ArtificialAnlys/status/2075639143372325205">Artificial Analysis</a>); <strong>CritPt</strong> gains over GPT-5.5 and beats Fable 5 by ~4 points (<a href="https://x.com/ArtificialAnlys/status/2075423964378366427">Artificial Analysis</a>); and strong results on <strong>WeirdML</strong> at lower cost (<a href="https://x.com/htihle/status/2075513299106426922">@htihle</a>). At the same time, users reported <strong>instruction-following issues</strong>, uneven token efficiency in practice, and some concern about <strong>jailbreakability / reward hacking</strong> (<a href="https://x.com/teortaxesTex/status/2075495527030964693">@teortaxesTex</a>, <a href="https://x.com/Mononofu/status/2075414796426764507">@Mononofu</a>, <a href="https://x.com/kimmonismus/status/2075693686604619948">@kimmonismus</a>).</p></li></ul><p><strong>Parallel-agent workflows, computer use, and the &#8220;harness is the product&#8221; theme</strong></p><ul><li><p><strong>GPT-5.6&#8217;s biggest perceived leap may be orchestration and computer use rather than pure chat quality</strong>. Multiple users highlighted that Sol is unusually strong as a <strong>planner / verifier / orchestrator</strong>, often using subagents automatically and reacting more quickly to steering (<a href="https://x.com/omarsar0/status/2075611352878481577">@omarsar0</a>, <a href="https://x.com/Hangsiin/status/2075463886309126271">@Hangsiin</a>). OpenAI also showcased <strong>computer use with Sol Ultra</strong> and promoted ChatGPT Work as bringing agents to consumer/mobile scale (<a href="https://x.com/gdb/status/2075619497764151644">OpenAI demo via @gdb</a>, <a href="https://x.com/gdb/status/2075628596232884556">Work positioning</a>). Community reports described very high-throughput GUI automation and Blender workflows (<a href="https://x.com/mckbrando/status/2075442660047814761">@mckbrando</a>, <a href="https://x.com/kimmonismus/status/2075482486901969066">@kimmonismus</a>).</p></li><li><p><strong>A recurring operational issue is hidden subagent cost explosion</strong>: users found that spawned agents may inherit premium settings, draining quotas much faster than expected. One concrete claim was that <code>spawn_agent</code> doesn&#8217;t let users choose model/effort, so <strong>Sol Ultra spawns more Sol Ultra</strong> by default (<a href="https://x.com/evi77ain/status/2075445272013095033">@evi77ain</a>). This fits the broader pattern of people liking the capability jump but finding the cost model opaque.</p></li><li><p><strong>The broader systems trend is toward harness-centric competition</strong>. This came through in product commentary from Perplexity&#8217;s Arav Srinivas (&#8220;the real product is now the harness around it&#8221;), in LangChain&#8217;s launch framing around <strong>Deep Agents + Nemotron + OpenShell</strong>, and in a growing set of memory / orchestration tools like <strong>OpenWiki</strong> and <strong>OpenSWE</strong> (<a href="https://x.com/dee_bosa/status/2075597686464491874">@dee_bosa quoting Arav</a>, <a href="https://x.com/hwchase17/status/2075620940466315608">@hwchase17</a>, <a href="https://x.com/BraceSproul/status/2075596668612014107">OpenWiki proactive memory</a>, <a href="https://x.com/BraceSproul/status/2075610067878257072">OpenSWE adoption</a>). The meta-point: frontier model parity is tightening, so value is increasingly shifting to <strong>routing, memory, tool use, safety rails, and enterprise context</strong>.</p></li></ul><p><strong>Meta&#8217;s Muse Spark 1.1 and the widening frontier of &#8220;good enough, fast, cheap&#8221; models</strong></p><ul><li><p><strong>Muse Spark 1.1 was the other major model story of the day</strong>, with many practitioners calling it the most surprising release of the week. Reports consistently emphasized <strong>strong UI/frontend generation, fast responses, and unusually aggressive pricing</strong>, often framing it as near-frontier quality for a large subset of coding/product tasks (<a href="https://x.com/alexandr_wang/status/2075652012608467385">@alexandr_wang</a>, <a href="https://x.com/rowancheung/status/2075634108324089943">@rowancheung</a>, <a href="https://x.com/kimmonismus/status/2075525943729275313">@kimmonismus</a>).</p></li><li><p><strong>Benchmarking suggests a real step up, but not outright frontier leadership</strong>. Artificial Analysis scored Muse Spark 1.1 at <strong>51</strong> on its Intelligence Index, up <strong>8 points</strong> from 1.0, roughly tied with <strong>GLM-5.2 / GPT-5.4 / GPT-5.6 Luna</strong> and behind <strong>Grok 4.5 / GPT-5.6 Sol / Claude Fable 5</strong>. Notable details: <strong>1M context</strong>, median speed ~<strong>114 tok/s</strong>, pricing <strong>$1.25 / $4.25 per 1M</strong> input/output tokens, and strong token efficiency (<a href="https://x.com/ArtificialAnlys/status/2075677416295739660">Artificial Analysis</a>). Arena also placed it <strong>#9 on Code Arena: Frontend</strong> with strong gains in instruction-following and longer-query categories (<a href="https://x.com/arena/status/2075642304501784698">Arena</a>).</p></li><li><p><strong>The strategic implication many drew</strong>: Meta&#8217;s compute-heavy bet is starting to show up as <strong>cost-effective inference products</strong>, not just talent headlines. Several commentators argued this materially raises competitive pressure on OpenAI/Anthropic, especially if Meta improves distribution and API ergonomics (<a href="https://x.com/scaling01/status/2075612353056342391">@scaling01 asking for OpenRouter</a>, <a href="https://x.com/alexandr_wang/status/2075680437620646370">@alexandr_wang</a>, <a href="https://x.com/mweinbach/status/2075600689200279747">@mweinbach</a>).</p></li></ul><p><strong>Open models, infra, and efficiency work</strong></p><ul><li><p><strong>Open-model tooling kept shipping despite the closed-model attention vacuum</strong>. Unsloth released <strong>Qwen3.6 NVFP4 quants</strong> with claims of <strong>2.5&#215; faster</strong> inference, including <strong>27B on 24GB VRAM</strong> and a <strong>35B-A3B</strong> variant hitting <strong>17,561 tok/s on B200</strong> (<a href="https://x.com/UnslothAI/status/2075566124687892597">Unsloth</a>, <a href="https://x.com/danielhanchen/status/2075567076002185525">technical details from @danielhanchen</a>). QuixiAI reported <strong>Qwen3.6-35B-A3B-NVFP4</strong> on dual B60 at <strong>65 tok/s</strong> and <strong>128k context</strong> (<a href="https://x.com/QuixiAI/status/2075418782470643958">QuixiAI</a>).</p></li><li><p><strong>Inference optimization remains a major live research area</strong>. Cohere open-sourced <strong>Hardware-aware Dynamic Speculative Decoding</strong> in vLLM, addressing the familiar issue where speculative decoding helps at low batch sizes but hurts at high ones (<a href="https://x.com/EkagraRanjan/status/2075640096829612416">Cohere/vLLM</a>, <a href="https://x.com/vllm_project/status/2075698626140295378">vLLM commentary</a>). Google/Hugging Face&#8217;s <strong>Gemma challenge</strong> reported up to <strong>5&#215; faster</strong> single-A10G inference, with <strong>315 TPS lossless</strong> and <strong>491.8 TPS</strong> fastest overall (<a href="https://x.com/googlegemma/status/2075611948985835877">Gemma</a>).</p></li><li><p><strong>Agent evaluation / self-improvement work is getting more concrete</strong>: &#8220;<strong>LLM-as-a-Verifier</strong>&#8221; reported SOTA on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench using repeated sampling plus score-logprob ranking (<a href="https://x.com/Azaliamirh/status/2075583355895058751">paper thread</a>); Meta researchers proposed an explicit memory agent to combat <strong>behavioral state decay</strong> in long-horizon agents (<a href="https://x.com/omarsar0/status/2075603504543269136">summary</a>).</p></li></ul><p><strong>Science, math, health, and modality-specific systems</strong></p><ul><li><p><strong>Math/science capability claims escalated sharply</strong>. OpenAI staff and community members circulated examples of <strong>GPT-5.6 Sol Ultra</strong> producing a claimed proof of the <strong>Cycle Double Cover Conjecture</strong> using <strong>64 subagents in under an hour</strong> (<a href="https://x.com/__eknight__/status/2075643450196971805">claim from @</a><strong><a href="https://x.com/__eknight__/status/2075643450196971805">eknight</a></strong>, <a href="https://x.com/gdb/status/2075670151702430044">amplified by @gdb</a>). Separately, Bubeck noted a single-person <strong>1M-line Lean formalization</strong> effort with GPT-5.6 (<a href="https://x.com/SebastienBubeck/status/2075407986772861047">@SebastienBubeck</a>). These are still claims pending external scrutiny, but they indicate where labs want the narrative to go: <strong>parallelized research agents as a scientific compute primitive</strong>.</p></li><li><p><strong>Health is becoming a first-class benchmark and product vertical</strong>. OpenAI said GPT-5.6 is a major step forward for <strong>health intelligence</strong>, highlighting that <strong>Luna at lowest effort beats GPT-5.5 at highest effort while costing 25&#215; less</strong> (<a href="https://x.com/OpenAI/status/2075686461693898868">OpenAI</a>). Karan Singhal added that, in blinded physician comparisons over <strong>20,000 axis ratings</strong>, physicians found <strong>fewer flaws in GPT-5.6 responses than physician-written responses</strong> across a hard task set (<a href="https://x.com/thekaransinghal/status/2075689779937833302">details</a>).</p></li><li><p><strong>Audio/music and creative tooling also moved</strong>: Kyutai + Mirelo released <strong>MuScriptor</strong>, an open model for <strong>multi-instrument audio-to-MIDI transcription from full mixes</strong>, not stems (<a href="https://x.com/MireloAI/status/2075536492177354771">MireloAI</a>, <a href="https://x.com/kyutai_labs/status/2075540047613276197">Kyutai</a>). Sakana&#8217;s new Picbreeder-style work explored <strong>open-ended creativity with VLM agents</strong>, concluding that diverse agent populations help but still fall short of human open-ended exploration (<a href="https://x.com/SakanaAILabs/status/2075580810330267844">Sakana</a>).</p></li></ul><p><strong>Security, safety, and policy frictions</strong></p><ul><li><p><strong>Security concerns rose alongside capability gains</strong>. OpenAI moved its <strong>Bio Bug Bounty</strong> into a private ongoing program and <strong>doubled rewards to $50K</strong>, specifically seeking universal jailbreaks against predefined biosafety challenges (<a href="https://x.com/OpenAI/status/2075647722766614733">OpenAI</a>). Separately, OpenAI tightened access requirements for its most cyber-capable models, requiring <strong>hardware security keys</strong> for Trusted Access for Cyber members starting Sept. 1 (<a href="https://x.com/cryps1s/status/2075639162120900766">@cryps1s</a>).</p></li><li><p><strong>Evidence of misuse remains salient</strong>: a new study reported <strong>Boko Haram</strong> members using frontier chatbots for bomb-making and related tactical queries (<a href="https://x.com/AntoniaJuelich/status/2075590815083028989">@AntoniaJuelich</a>). That thread sat uncomfortably next to ongoing online discussion that GPT-5.6 may be relatively easy to jailbreak or reward-hack in some settings (<a href="https://x.com/Mononofu/status/2075414796426764507">@Mononofu</a>).</p></li><li><p><strong>Policy discourse remains polarized and speculative</strong>. The &#8220;AI 2040 / Plan A&#8221; transparency-and-governance scenario drew both support and ridicule, with Ajeya Cotra emphasizing the centrality of <strong>total research transparency</strong> while critics questioned feasibility and assumptions about superintelligence/governance capacity (<a href="https://x.com/ajeya_cotra/status/2075583823434371250">@ajeya_cotra</a>, <a href="https://x.com/binarybits/status/2075660927001608431">@binarybits</a>, <a href="https://x.com/banteg/status/2075512151783972925">@banteg satire</a>).</p></li></ul><p><strong>Top tweets (by engagement)</strong></p><ul><li><p><strong>OpenAI launch and rollback management</strong>: OpenAI&#8217;s product lead acknowledged launch confusion, promised UI fixes, and reset usage twice while clarifying that <strong>Codex is here to stay</strong> (<a href="https://x.com/thsottiaux/status/2075641131002700120">full thread</a>).</p></li><li><p><strong>Claude Code desktop browser</strong>: Anthropic shipped an <strong>in-app browser</strong> for Claude Code desktop so Claude can browse docs/sites inside the app (<a href="https://x.com/ClaudeDevs/status/2075635283211772279">@ClaudeDevs</a>).</p></li><li><p><strong>OpenAI org update</strong>: Fidji Simo announced she is leaving her full-time role at OpenAI and becoming a <strong>part-time advisor</strong>, citing the need to focus on recovery from chronic illness while continuing work related to AI and health (<a href="https://x.com/fidjissimo/status/2075353170927304861">@fidjissimo</a>).</p></li><li><p><strong>Perplexity harness expansion</strong>: Perplexity added <strong>Grok 4.5</strong> as an orchestrator in Computer after internal evals showed strong WANDR performance at roughly half the cost of Opus 4.8 (<a href="https://x.com/perplexity_ai/status/2075660058625790159">Perplexity</a>).</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. GLM-5.2 Local Inference and Security Scrutiny</strong></h3><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-not-much-happened-today-f5c">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[[AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp]]></title><description><![CDATA[A big day for OpenAI.]]></description><link>https://www.latent.space/p/ainews-openai-launches-gpt-56-solterraluna</link><guid isPermaLink="false">https://www.latent.space/p/ainews-openai-launches-gpt-56-solterraluna</guid><dc:creator><![CDATA[Latent.Space]]></dc:creator><pubDate>Fri, 10 Jul 2026 06:19:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/-MPGU2a67Ls" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>On any other day, the launch of a surprisingly good/competitive <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Muse Spark 1.1</a> from Meta Superintelligence Labs, including, for the first time, in the <a href="https://developer.meta.com/ai/resources/blog/build-with-muse-spark/">Meta Model API</a> (signaling high confidence for broad usage and third party testing which <a href="https://x.com/alexandr_wang/status/2074687661428572403">is bearing out in their sister models</a>), would deserve title story status, but they had the misfortune of going up against a mainline frontier model launch:</p><div id="youtube2--MPGU2a67Ls" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;-MPGU2a67Ls&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/-MPGU2a67Ls?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>As <a href="https://openai.com/index/previewing-gpt-5-6-sol/">previewed a couple weeks ago</a> before government approval, 5.6 comes in three new sizes, Sol, Terra and Luna, corresponding to the sizes of Sun, Earth and Moon, as an alternative to the more literary sizing of Claude variants, and a new <code>ultra</code> effort level, <em>&#8220;our highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster&#8221;:</em></p><blockquote><p><code>max</code><em> gives GPT&#8209;5.6 even more time than </em><code>xhigh</code><em> to reason and explore alternatives, run checks, and revise its approach. ultra goes further by <strong>coordinating four agents in parallel by default</strong>, trading higher token use for stronger results and faster time-to-result on demanding tasks.</em> </p></blockquote><p>On multiple benchmarks (not just the ones featured here), 5.6 both achieves higher performance at lower cost than Fable or Opus.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S2WI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S2WI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 424w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 848w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 1272w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S2WI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png" width="1456" height="1393" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1393,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:136495,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206398209?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!S2WI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 424w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 848w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 1272w, https://substackcdn.com/image/fetch/$s_!S2WI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d7bcf2a-7c60-4e1d-9aaf-db0b03cc4801_1470x1406.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><em>&#8220;Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. It also sets new state-of-the-art results on Terminal&#8209;Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases.&#8221;</em></p></blockquote><p>There are also harder-to-benchmark improvements in computer use, presentation/document generation, and scientific research that should nevertheless be taken very seriously.</p><p></p><p>As we <a href="https://www.latent.space/p/ainews-gpt-55-and-openai-codex-superapp?utm_source=publication-search">predicted in April</a>, the newly launched <a href="https://x.com/OpenAI/status/2075274271845404744?s=20">ChatGPT Work</a> and Codex desktop app update today is probably the penultimate step for OpenAI&#8217;s superapp strategy (the last open question is what happens to the agentic browser&#8230;.)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SAjG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SAjG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 424w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 848w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 1272w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SAjG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png" width="846" height="1090" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1090,&quot;width&quot;:846,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:295153,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206398209?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SAjG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 424w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 848w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 1272w, https://substackcdn.com/image/fetch/$s_!SAjG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1521892-45ba-4676-a486-65663d1a6bb9_846x1090.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p></p><p></p><p></p><p></p><blockquote><p>AI News for 7/08/2026-7/09/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>OpenAI launched a new three-model GPT&#8209;5.6 family and simultaneously expanded the product stack around it.</strong></p><ul><li><p>OpenAI announced <strong>GPT&#8209;5.6 Sol, Terra, and Luna</strong> rolling out across <strong>ChatGPT, Codex, and the API</strong> via <a href="https://x.com/OpenAI/status/2075271421149020426">@OpenAI</a> and <a href="https://x.com/OpenAIDevs/status/2075273992609599834">@OpenAIDevs</a></p></li><li><p>In ChatGPT, <strong>Plus, Pro, Business, and Enterprise</strong> users get access to <strong>GPT&#8209;5.6 Sol</strong> through medium+ effort settings, while <strong>Pro and Enterprise</strong> can select <strong>GPT&#8209;5.6 Pro</strong> for highest-quality results on complex tasks, per <a href="https://x.com/OpenAI/status/2075271435573244008">@OpenAI</a></p></li><li><p>API pricing introduced a tiered lineup: <strong>Sol $5 / $30 per million input/output tokens</strong>, <strong>Terra $2.5 / $15</strong>, <strong>Luna $1 / $6</strong>, with <strong>cache-write pricing</strong> added for the first time and <strong>90% cache-read discount</strong> retained, according to <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a></p></li><li><p>OpenAI framed the family around a price-performance ladder: <strong>Sol = flagship/highest ceiling</strong>, <strong>Terra = GPT&#8209;5.5-like capability at lower cost</strong>, <strong>Luna = fastest/cheapest high-volume option</strong>, via <a href="https://x.com/OpenAIDevs/status/2075286157186003348">@OpenAIDevs</a></p></li><li><p>The launch bundled major app-layer changes: <strong>ChatGPT Work</strong>, a new <strong>desktop app merging Codex + ChatGPT</strong>, <strong>Sites</strong> beta, <strong>programmatic tool calling</strong>, and <strong>multi-agent beta</strong> in the Responses API, via <a href="https://x.com/OpenAI/status/2075274271845404744">@OpenAI</a>, <a href="https://x.com/OpenAIDevs/status/2075275868268789885">@OpenAIDevs</a>, and <a href="https://x.com/OpenAIDevs/status/2075274093327470923">@OpenAIDevs</a></p></li></ul><h2><strong>Official claims and benchmark results</strong></h2><p><strong>OpenAI&#8217;s official message emphasized strong agentic/coding performance, better artifact quality, and improved economics.</strong></p><ul><li><p>Sam Altman called it &#8220;<strong>obviously the best model we have ever produced</strong>&#8221; in the launch post, linking the release blog, via <a href="https://x.com/sama/status/2075266471316615436">@sama</a></p></li><li><p>Altman also highlighted enterprise economics: &#8220;<strong>5.6 sol is a huge step forward for dollars-per-task</strong>,&#8221; via <a href="https://x.com/sama/status/2075267201058426944">@sama</a></p></li><li><p>Greg Brockman said the goal is &#8220;<strong>the best price for any level of target performance</strong>&#8221; and the highest possible ceiling, via <a href="https://x.com/gdb/status/2075271293474353553">@gdb</a></p></li><li><p>OpenAI claimed <strong>GPT&#8209;5.6 Sol sets a new high of 53.6 on Agents&#8217; Last Exam</strong>, beating <strong>Claude Fable 5 adaptive by 13.1 points</strong>; at medium reasoning it beats Fable by <strong>11.4 points at roughly one-quarter the estimated cost</strong>, while <strong>Terra and Luna also outperform Fable at around one-sixteenth the cost</strong>, via <a href="https://x.com/OpenAI/status/2075271423992680532">@OpenAI</a></p></li><li><p>OpenAI said GPT&#8209;5.6 improves <strong>artifact quality across presentations, documents, and spreadsheets</strong>, with outputs exportable into existing enterprise tools, via <a href="https://x.com/OpenAI/status/2075271432041545782">@OpenAI</a></p></li><li><p>OpenAI positioned GPT&#8209;5.6 as state of the art for <strong>reasoning through complex tasks</strong> and for producing materials matched to templates, reference files, and preferred style inside <strong>ChatGPT Work</strong>, via <a href="https://x.com/OpenAI/status/2075274275104399670">@OpenAI</a></p></li><li><p>OpenAI also said GPT&#8209;5.6 is its <strong>most capable model yet on cyber and bio-related tasks</strong>, with some API calls potentially blocked or paused for extra safety review in dual-use areas, via <a href="https://x.com/OpenAIDevs/status/2075274080740380829">@OpenAIDevs</a></p></li><li><p>OpenAI highlighted better <strong>Computer Use</strong> performance: faster, more token-efficient, support for <strong>batching and parallel operations</strong> across multi-step tasks, plus picture-in-picture supervision, via <a href="https://x.com/OpenAIDevs/status/2075276074980884862">@OpenAIDevs</a></p></li></ul><h2><strong>Independent evaluations and third-party measurements</strong></h2><p><strong>Independent evals broadly placed Sol near or at the frontier, especially on coding-agent workloads, while also surfacing caveats.</strong></p><ul><li><p><a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a> reported <strong>GPT&#8209;5.6 Sol (max)</strong> scores <strong>59</strong> on its Intelligence Index, <strong>1 point below Claude Fable 5 (max)</strong>, at <strong>about one-third of Fable&#8217;s cost per task</strong></p></li><li><p>On the same analysis, <strong>Terra</strong> and <strong>Luna</strong> score <strong>55</strong> and <strong>51</strong> on the Intelligence Index, with <strong>~50%</strong> and <strong>~80%</strong> lower cost per task than Sol, respectively, via <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a></p></li><li><p>Artificial Analysis said <strong>Sol leads the Coding Agent Index at 80</strong>, ahead of Fable 5 and Opus 4.8, and is also cheaper per task than both on their harnesses, via <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a></p></li><li><p>It also noted <strong>Sol defines a new Pareto frontier of intelligence vs output tokens</strong>, while <strong>Terra and Luna are not on that frontier</strong>, via <a href="https://x.com/ArtificialAnlys/status/2075268984539410521">@ArtificialAnlys</a></p></li><li><p>Artificial Analysis found <strong>minor improvement over GPT&#8209;5.5 in AA&#8209;Omniscience</strong> but with a <strong>higher hallucination rate</strong> than GPT&#8209;5.5 max, via <a href="https://x.com/ArtificialAnlys/status/2075268990004605023">@ArtificialAnlys</a></p></li><li><p>It reported <strong>similar GDPval-AA v2 performance to Claude Fable 5</strong>, suggesting comparable ability on economically valuable tasks, via <a href="https://x.com/ArtificialAnlys/status/2075268987550932998">@ArtificialAnlys</a></p></li><li><p><a href="https://x.com/ValsAI/status/2075270642359029972">@ValsAI</a> ranked GPT&#8209;5.6 <strong>#2 on Vals Index and Vals Multimodal Index</strong>, saying Fable 5 remains ahead on several benchmarks but GPT&#8209;5.6 is &#8220;clearly in the same class&#8221;</p></li><li><p>Vals also said <strong>Sol is #1 on CyberBench and Excel Modeling Benchmark</strong>, and #1 on <strong>Legal Research Bench, ProofBench, SWE-bench, and Terminal-Bench 2.1</strong>, adding that Fable had a nearly <strong>100% refusal rate on CyberBench</strong>, via <a href="https://x.com/ValsAI/status/2075270644711997581">@ValsAI</a></p></li><li><p><a href="https://x.com/arcprize/status/2075270869992264003">@arcprize</a> said <strong>GPT&#8209;5.6 Sol scores 7.8% on ARC&#8209;AGI&#8209;3</strong> and is the <strong>first verified frontier model to ever beat an ARC&#8209;AGI&#8209;3 game</strong></p></li><li><p><a href="https://x.com/GregKamradt/status/2075274981794300113">@GregKamradt</a> noted <strong>92.5% on ARC&#8209;AGI&#8209;2</strong>, calling it SOTA while costing <strong>an order of magnitude less</strong> than GPT&#8209;5.5 Pro three months earlier</p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2075423964378366427">@ArtificialAnlys</a> later reported <strong>GPT&#8209;5.6 Sol (max) leads CritPt</strong>, a benchmark of unpublished research-level physics problems, by roughly <strong>4 points over Claude Fable 5</strong></p></li><li><p><a href="https://x.com/llama_index/status/2075351095258296378">@llama_index</a> said day-0 ParseBench results show GPT&#8209;5.6 continues to do well on <strong>text and tables</strong> but still struggles on <strong>charts and layout</strong>, and that <strong>Luna is ~6&#215; cheaper than Sol with only minor degradations</strong></p></li><li><p><a href="https://x.com/jerryjliu0/status/2075356305099800717">@jerryjliu0</a> similarly said ParseBench shows <strong>no high-level change versus GPT&#8209;5.5</strong> on tables/text/charts/layout, stressing persistent weakness on <strong>complex text layouts, chart transcription, and source-element bounding boxes</strong></p></li></ul><h2><strong>Technical details</strong></h2><p><strong>The technical story of GPT&#8209;5.6 is as much about inference orchestration and token efficiency as raw capability.</strong></p><ul><li><p>OpenAI shipped <strong>three model tiers</strong> with multiple <strong>reasoning effort levels</strong>; users discussed <strong>Light, Medium, High, Extra High, Ultra</strong>, leading to a large configuration matrix, via <a href="https://x.com/rasbt/status/2075369179817902176">@rasbt</a></p></li><li><p>OpenAI added <strong>Programmatic Tool Calling</strong> in the Responses API and <strong>Multi-agent beta</strong>, indicating more explicit support for orchestrated tool use and agent decomposition, via <a href="https://x.com/OpenAIDevs/status/2075274093327470923">@OpenAIDevs</a></p></li><li><p>OpenAI&#8217;s app layer now uses <strong>Codex as the core</strong> of the new Work product, per <a href="https://x.com/sama/status/2075293792048136572">@sama</a> and <a href="https://x.com/gdb/status/2075276416686723110">@gdb</a></p></li><li><p>Several posts stress <strong>parallel agents/subagents</strong> as a major capability lever; <a href="https://x.com/aidan_mclau/status/2075337767949865464">@aidan_mclau</a> explicitly mentions users can increase the number of <strong>5.6 subagents</strong></p></li><li><p><a href="https://x.com/LiorOnAI/status/2075277748394967122">@LiorOnAI</a> summarized likely drivers as <strong>adaptive reasoning</strong>, <strong>parallel agents</strong>, <strong>programmatic tool use</strong>, and <strong>higher token efficiency</strong></p></li><li><p>Artificial Analysis reported <strong>Sol max uses ~15k output tokens per Intelligence Index task vs 16k for GPT&#8209;5.5</strong>, and fewer than Opus 4.8, GLM&#8209;5.2, and Gemini 3.5 Flash at comparable intelligence, via <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a></p></li><li><p><a href="https://x.com/OpenRouter/status/2075271807855452196">@OpenRouter</a> said early testing found the 5.6 models <strong>more token efficient</strong>, lowering both cost and time-to-task completion</p></li><li><p>The desktop/app layer brought a <strong>Chrome extension</strong>, <strong>revamped in-app browser</strong>, <strong>authenticated sites</strong>, <strong>persistent multi-tab sessions</strong>, <strong>file downloads</strong>, and tighter cross-device handoffs, via <a href="https://x.com/OpenAIDevs/status/2075275868268789885">@OpenAIDevs</a>, <a href="https://x.com/OpenAIDevs/status/2075276009902112976">@OpenAIDevs</a>, and <a href="https://x.com/OpenAIDevs/status/2075292716737736919">@OpenAIDevs</a></p></li><li><p><strong>Sites</strong> entered beta for paid users, offering hosting, storage, and optional auth for GPT-built apps, via <a href="https://x.com/OpenAIDevs/status/2075275892591591469">@OpenAIDevs</a> and <a href="https://x.com/OpenAIDevs/status/2075337081304522853">@OpenAIDevs</a></p></li></ul><h2><strong>The &#8220;Sol autonomously post-trained Luna&#8221; claim</strong></h2><p><strong>This was the most provocative technical claim around the launch, but its interpretation became contested almost immediately.</strong></p><ul><li><p>Multiple accounts amplified the statement that <strong>OpenAI says GPT&#8209;5.6 Sol autonomously post-trained GPT&#8209;5.6 Luna</strong>, via <a href="https://x.com/scaling01/status/2075269113488789984">@scaling01</a>, <a href="https://x.com/tejalpatwardhan/status/2075272564629451110">@tejalpatwardhan</a>, and <a href="https://x.com/dejavucoder/status/2075270116909232129">@dejavucoder</a></p></li><li><p>The claim fueled RSI/autoresearch speculation; <a href="https://x.com/tenobrus/status/2075282678652522712">@tenobrus</a> said if true as stated, it would be a &#8220;pretty large update&#8221; for automated researcher timelines</p></li><li><p><a href="https://x.com/eliebakouch/status/2075281402807844872">@eliebakouch</a> framed it as OpenAI asking Sol to post-train Luna &#8220;with <strong>100k GPUs</strong>&#8221; for an experiment</p></li><li><p><a href="https://x.com/gdb/status/2075363531042726216">@gdb</a> said the implication is easy to overlook for accelerating engineering workflows, reinforcing that OpenAI wants this read as more than a marketing flourish</p></li><li><p>But skeptical clarifications emerged quickly: <a href="https://x.com/nikolaj2030/status/2075297831376793764">@nikolaj2030</a> asked whether this actually meant Sol completed a <strong>small controlled post-training task</strong>&#8212;modifying a config, editing a scheduler file, and launching a run&#8212;rather than end-to-end real-world post-training of Luna</p></li><li><p><a href="https://x.com/nrehiew_/status/2075316190386462888">@nrehiew_</a> interpreted the screenshot similarly: Sol could go from high-level ideas to <strong>editing configs and launching experiments</strong>, not fully owning Luna&#8217;s end-to-end post-training</p></li><li><p><a href="https://x.com/scaling01/status/2075354327791587467">@scaling01</a> argued that what&#8217;s probably happening is a model implementing <strong>LLM-as-a-judge graders</strong>, reward-shaping logic, or small training configs on top of existing OpenAI RL infrastructure&#8212;not autonomous end-to-end research or training systems</p></li><li><p><a href="https://x.com/scaling01/status/2075359429717836251">@scaling01</a> explicitly said we should distance these statements from <strong>literal autonomous end-to-end post-training or research</strong>, which models still cannot do</p></li><li><p>Counterbalancing that skepticism, <a href="https://x.com/aidan_mclau/status/2075328409400738229">@aidan_mclau</a> said it is routine for him to have <strong>5.6 e2e do an entire RL run</strong>, suggesting meaningful internal workflow automation even if not self-sufficient research</p></li><li><p>The consensus across technical observers was not that Sol independently invented and trained Luna, but that GPT&#8209;5.6 may now be capable of <strong>executing meaningful chunks of model-improvement workflows inside mature internal infrastructure</strong></p></li></ul><h2><strong>Internal productivity and recursive improvement signals</strong></h2><p><strong>OpenAI also used internal-usage data to argue that GPT&#8209;5.6 materially changes researcher throughput.</strong></p><ul><li><p><a href="https://x.com/scaling01/status/2075269455781703850">@scaling01</a> highlighted an OpenAI claim that it <strong>doubled experiment throughput per researcher</strong> since the start of the year</p></li><li><p><a href="https://x.com/eliebakouch/status/2075273299148341327">@eliebakouch</a> quoted OpenAI saying average daily output tokens per active researcher were <strong>more than twice the highest level observed for GPT&#8209;5.5</strong> during internal testing</p></li><li><p>Another OpenAI stat, relayed by <a href="https://x.com/eliebakouch/status/2075273992185782661">@eliebakouch</a>, said over six months the share of research compute devoted to <strong>internal coding inference grew 100-fold</strong>, while <strong>internal agentic token usage increased ~22-fold</strong></p></li><li><p><a href="https://x.com/FakePsyho/status/2075291659814781370">@FakePsyho</a> linked these developments to OpenAI&#8217;s performance in top programming contests, describing systems close to GPT&#8209;5.6 plus custom harnesses as decisively beating elite human competitors</p></li><li><p>This fed broader RSI/autoresearch discussion, especially from people who see long-horizon coding and heuristic optimization as proxies for model-improvement capability</p></li></ul><h2><strong>Product implications: ChatGPT Work, Codex merge, desktop, and Sites</strong></h2><p><strong>The model launch doubled as a product strategy reset: OpenAI is pushing from &#8220;chatbot&#8221; to &#8220;work OS.&#8221;</strong></p><ul><li><p>OpenAI launched <strong>ChatGPT Work</strong>, an agent powered by <strong>Codex + GPT&#8209;5.6</strong> that can act across apps and files, stay on tasks for hours, and turn a goal into finished work, via <a href="https://x.com/OpenAI/status/2075274271845404744">@OpenAI</a></p></li><li><p>Work can ingest context from <strong>docs, Slack, Notion, Microsoft 365, and Google Drive</strong> and produce <strong>decks, docs, spreadsheets, dashboards, visualizations, and interactive explanations</strong>, summarized by <a href="https://x.com/kimmonismus/status/2075271465964798147">@kimmonismus</a></p></li><li><p>The <strong>Codex app merged into the new ChatGPT desktop app</strong>, confirmed by <a href="https://x.com/avstorm/status/2075266403297362364">@avstorm</a> and <a href="https://x.com/OpenAIDevs/status/2075275880704995342">@OpenAIDevs</a></p></li><li><p>Developers now get <strong>inline diff editing</strong>, <strong>PR review side panel</strong>, better <strong>SSH video rendering</strong>, and stronger <strong>computer use</strong>, via <a href="https://x.com/romainhuet/status/2075286364476850430">@romainhuet</a> and <a href="https://x.com/reach_vb/status/2075280626362560805">@reach_vb</a></p></li><li><p><strong>Sites</strong> lets users turn work into shareable hosted apps/websites from ChatGPT, via <a href="https://x.com/OpenAIDevs/status/2075275892591591469">@OpenAIDevs</a> and <a href="https://x.com/simpsoka/status/2075278935366287842">@simpsoka</a></p></li><li><p><a href="https://x.com/OpenAI/status/2075310019185389913">@OpenAI</a>, <a href="https://x.com/OpenAI/status/2075310020653351324">@OpenAI</a>, and <a href="https://x.com/OpenAI/status/2075310022121472399">@OpenAI</a> marketed GPT&#8209;5.6 through case studies: a <strong>broccoli farmer</strong>, a <strong>mathematician</strong>, and a <strong>family cereal business</strong></p></li><li><p>This product reframing was read by some as OpenAI&#8217;s answer to Anthropic&#8217;s Cowork / Claude Code stack, via <a href="https://x.com/jerryjliu0/status/2075295459304710496">@jerryjliu0</a> and <a href="https://x.com/kimmonismus/status/2075280933452669000">@kimmonismus</a></p></li></ul><h2><strong>Facts vs opinions</strong></h2><p><strong>Facts / directly sourced claims</strong></p><ul><li><p>GPT&#8209;5.6 family names, rollout channels, and access tiers: <a href="https://x.com/OpenAI/status/2075271421149020426">@OpenAI</a>, <a href="https://x.com/OpenAI/status/2075271435573244008">@OpenAI</a>, <a href="https://x.com/OpenAIDevs/status/2075273992609599834">@OpenAIDevs</a></p></li><li><p>API prices and cache-write policy: <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a></p></li><li><p>OpenAI&#8217;s benchmark claims on Agents&#8217; Last Exam: <a href="https://x.com/OpenAI/status/2075271423992680532">@OpenAI</a></p></li><li><p>Artificial Analysis and Vals leaderboard placements: <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a>, <a href="https://x.com/ValsAI/status/2075270642359029972">@ValsAI</a></p></li><li><p>ARC&#8209;AGI&#8209;3 7.8% claim: <a href="https://x.com/arcprize/status/2075270869992264003">@arcprize</a></p></li><li><p>ParseBench caveats: <a href="https://x.com/llama_index/status/2075351095258296378">@llama_index</a>, <a href="https://x.com/jerryjliu0/status/2075356305099800717">@jerryjliu0</a></p></li><li><p>Safety testing finding jailbreaks on GPT&#8209;5.6 Sol: <a href="https://x.com/alxndrdavies/status/2075279477626564933">@alxndrdavies</a></p></li></ul><p><strong>Opinions / interpretation / hype</strong></p><ul><li><p>&#8220;Best model we have ever produced&#8221;: <a href="https://x.com/sama/status/2075266471316615436">@sama</a></p></li><li><p>&#8220;First time I&#8217;ve felt comfortable delegating the hardest problem out there&#8221;: <a href="https://x.com/reach_vb/status/2075269547439907269">@reach_vb</a></p></li><li><p>&#8220;Not enough people are emotionally prepared for GPT&#8209;6&#8221;: <a href="https://x.com/scaling01/status/2075276735650648258">@scaling01</a></p></li><li><p>&#8220;OpenAI is competing on cost curves, not benchmarks&#8221;: <a href="https://x.com/LiorOnAI/status/2075277748394967122">@LiorOnAI</a></p></li><li><p>&#8220;The engineers were allowed to cook&#8221;: <a href="https://x.com/TheHumanoidHub/status/2075272514755059773">@TheHumanoidHub</a></p></li><li><p>&#8220;Generational fumble&#8221; regarding Codex becoming ChatGPT Desktop: <a href="https://x.com/theo/status/2075312087723876556">@theo</a></p></li></ul><h2><strong>Different perspectives</strong></h2><p><strong>Supportive views</strong></p><ul><li><p>Many developers and evaluators saw GPT&#8209;5.6 as a meaningful frontier advance, especially in coding and knowledge work: <a href="https://x.com/gdb/status/2075270503405924466">@gdb</a>, <a href="https://x.com/AravSrinivas/status/2075270640177938547">@AravSrinivas</a>, <a href="https://x.com/OpenRouter/status/2075271807855452196">@OpenRouter</a>, <a href="https://x.com/Teknium/status/2075392507794624803">@Teknium</a></p></li><li><p>Several posts focused on <strong>cost efficiency</strong> as the real win, with Sol matching frontier peers while being materially cheaper: <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a>, <a href="https://x.com/omarsar0/status/2075270117131259925">@omarsar0</a>, <a href="https://x.com/cline/status/2075278343927365991">@cline</a></p></li><li><p>Others highlighted the <strong>agentic stack</strong>&#8212;Work, Codex, multi-agent, programmatic tools&#8212;as more strategically important than raw benchmark deltas: <a href="https://x.com/TheRundownAI/status/2075273458661949763">@TheRundownAI</a>, <a href="https://x.com/kimmonismus/status/2075271465964798147">@kimmonismus</a>, <a href="https://x.com/fidjissimo/status/2075305622120325363">@fidjissimo</a></p></li></ul><p><strong>Neutral / analytical views</strong></p><ul><li><p>Some analysts saw Sol as roughly <strong>same class as Fable</strong>, but not decisively ahead overall: <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a>, <a href="https://x.com/ValsAI/status/2075270642359029972">@ValsAI</a></p></li><li><p><a href="https://x.com/teortaxesTex/status/2075274583226069040">@teortaxesTex</a> argued the release may reflect OpenAI strong post-training recovering toward Anthropic despite a stronger Anthropic base model</p></li><li><p><a href="https://x.com/simonw/status/2075306164993315192">@simonw</a> pointed to notable API additions but also implied growing product complexity</p></li></ul><p><strong>Critical / skeptical views</strong></p><ul><li><p><a href="https://x.com/scaling01/status/2075268278105067566">@scaling01</a> asked whether <strong>GPT&#8209;5.6 Sol is worse at math</strong>, pushing back on the &#8220;everything got better&#8221; narrative</p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2075268990004605023">@ArtificialAnlys</a> found <strong>higher hallucination rate vs GPT&#8209;5.5</strong></p></li><li><p><a href="https://x.com/scaling01/status/2075279452494299273">@scaling01</a> criticized the ARC&#8209;AGI&#8209;3 scoring setup, saying Sol would score <strong>0% under official scoring methodology capped at $10k</strong> and objecting to use of a <strong>$25k</strong> budget</p></li><li><p><a href="https://x.com/Hangsiin/status/2075277820528607704">@Hangsiin</a> and <a href="https://x.com/Hangsiin/status/2075278682160275561">@Hangsiin</a> pointed to <strong>subscription/credit confusion</strong>, saying Sol costs more credits than GPT&#8209;5.5 while usage limits differ less than API pricing suggests</p></li><li><p><a href="https://x.com/QuinnyPig/status/2075334468462899442">@QuinnyPig</a> said OpenAI&#8217;s pricing/subscription strategy is confusing, particularly around future pricing jumps or inclusion terms</p></li><li><p><a href="https://x.com/rasbt/status/2075369179817902176">@rasbt</a> highlighted UX complexity: <strong>2 modes &#215; 3 models &#215; 5 effort levels = 30 configurations</strong></p></li><li><p><a href="https://x.com/MParakhin/status/2075361980446289925">@MParakhin</a> complained that <strong>GPT&#8209;5.6 Pro no longer has extended thinking</strong>, preferring an option to pay for much longer reasoning</p></li><li><p><a href="https://x.com/theo/status/2075312087723876556">@theo</a> and <a href="https://x.com/simonw/status/2075348941215006888">@simonw</a> criticized the growing app/mode fragmentation around ChatGPT, Codex, and Work</p></li></ul><h2><strong>Safety and security concerns</strong></h2><p><strong>The launch also surfaced one of the strongest public cyber-safety debates around a recent frontier model release.</strong></p><ul><li><p><a href="https://x.com/alxndrdavies/status/2075279477626564933">@alxndrdavies</a> from the AI Safety Institute said they found <strong>universal jailbreaks in all rounds of testing</strong> that enabled long-form agentic task completion in <strong>vulnerability discovery and exploit development</strong></p></li><li><p><a href="https://x.com/EthanJPerez/status/2075296476817985751">@EthanJPerez</a> called it &#8220;<strong>the highest stakes safety issue of any model release yet</strong>&#8221;</p></li><li><p><a href="https://x.com/yonashav/status/2075286161241612664">@yonashav</a> praised OpenAI for allowing third-party unreleased-model safety assessments to be published even when inconvenient</p></li><li><p><a href="https://x.com/Mononofu/status/2075414796426764507">@Mononofu</a> said ease of jailbreaking plus reward-hacking reports make them worried OpenAI may have rushed the release to keep pace with Fable</p></li><li><p>At the same time, OpenAI explicitly warned some cyber/bio requests may be paused or blocked mid-stream for additional review, via <a href="https://x.com/OpenAIDevs/status/2075274080740380829">@OpenAIDevs</a></p></li><li><p>This created a split narrative: strong cyber capability is treated as a product advantage by some evaluators, but as a serious deployment risk by safety researchers</p></li></ul><h2><strong>Context</strong></h2><p><strong>Why this matters goes beyond a single model benchmark win.</strong></p><ul><li><p>The launch happened amid a compressed week of frontier competition that also included new releases from <strong>Meta Muse Spark 1.1</strong> and <strong>Grok 4.5</strong>, leading multiple observers to describe the frontier as newly crowded: <a href="https://x.com/matanSF/status/2075276339607654802">@matanSF</a>, <a href="https://x.com/kimmonismus/status/2075322537592922345">@kimmonismus</a></p></li><li><p>OpenAI&#8217;s differentiation is increasingly framed less as &#8220;best raw benchmark score&#8221; and more as <strong>cost-efficient agentic work</strong>, consistent with posts from <a href="https://x.com/sama/status/2075267201058426944">@sama</a>, <a href="https://x.com/ArtificialAnlys/status/2075268970492657905">@ArtificialAnlys</a>, and <a href="https://x.com/LiorOnAI/status/2075277748394967122">@LiorOnAI</a></p></li><li><p>The product bundling suggests OpenAI is moving from a model vendor to a <strong>full-stack work platform</strong>, with its own browser, connectors, orchestration primitives, hosted app deployment, and desktop runtime</p></li><li><p>The strongest forward-looking signal may be the internal claim that researchers already use these systems to materially increase output and automate chunks of RL/post-training workflows, even if public discussion often overstates that as &#8220;the model trained itself&#8221;</p></li><li><p>The launch also sharpens a recurring engineering question raised by many tweets: whether the frontier is now bottlenecked less by a single monolithic model and more by <strong>orchestration quality, tool APIs, subagents, evaluation harnesses, and economics</strong></p></li></ul><p><strong>Frontier models and evaluations</strong></p><ul><li><p><strong>Meta launched Muse Spark 1.1</strong> and the <strong>Meta Model API</strong> in public preview, positioning it as a strong <strong>agentic, coding, multimodal, and computer-use</strong> model. Official posts came from <a href="https://x.com/finkd/status/2075218444056707458">@finkd</a>, <a href="https://x.com/alexandr_wang/status/2075218936266998230">@alexandr_wang</a>, <a href="https://x.com/shengjia_zhao/status/2075220782465290620">@shengjia_zhao</a>, <a href="https://x.com/ren_hongyu/status/2075224643829711101">@ren_hongyu</a>, and <a href="https://x.com/MetaforDevs/status/2075268072022401526">@OpenAIDevs</a></p></li><li><p>Key technical details repeatedly cited: <strong>1M-token context window</strong>, <strong>video understanding</strong>, multimodal reasoning, and API availability, with <a href="https://x.com/altryne/status/2075237837033889911">@altryne</a> and <a href="https://x.com/xinyun_chen_/status/2075276047495659656">@xinyun_chen_</a> among those emphasizing long-horizon agentic gains</p></li><li><p>Benchmark claims around Muse Spark 1.1 included competitiveness with <strong>GPT&#8209;5.5</strong> and <strong>Opus 4.8</strong> on agentic evals, strong performance on <strong>Harvey&#8217;s Legal Bench, TaxEval, MedScribe</strong>, and some out-of-distribution evals over <strong>Opus 4.8</strong> and <strong>Grok 4.5</strong>, via <a href="https://x.com/alexandr_wang/status/2075233663323947120">@alexandr_wang</a>, <a href="https://x.com/alexandr_wang/status/2075275671815999956">@alexandr_wang</a>, <a href="https://x.com/_jasonwei/status/2075265159430623334">@_jasonwei</a>, and <a href="https://x.com/cline/status/2075271057326719152">@cline</a></p></li><li><p>External reaction ranged from surprise and enthusiasm&#8212;e.g. <a href="https://x.com/kimmonismus/status/2075232528726708245">@kimmonismus</a>, <a href="https://x.com/preston_ojb/status/2075229604244271470">@preston_ojb</a>, <a href="https://x.com/0interestrates/status/2075330028729143634">@0interestrates</a>&#8212;to practical integration pushes from <a href="https://x.com/cline/status/2075271057326719152">@cline</a></p></li><li><p><strong>Grok 4.5</strong> continued to draw benchmark discussion: <a href="https://x.com/arena/status/2075301317560742373">@arena</a> said it reached <strong>#3 in Code Arena: Frontend</strong>, while <a href="https://x.com/alexgshaw/status/2075273675331580218">@alexgshaw</a> discussed <strong>Terminal-Bench 2.1</strong> reward-hacking caveats. Several posters argued Grok now belongs in the frontier set, including <a href="https://x.com/teortaxesTex/status/2075347335412953265">@teortaxesTex</a></p></li></ul><p><strong>Agents, orchestration, and developer tooling</strong></p><ul><li><p>Multiple posts reinforced that <strong>harness/orchestration quality</strong> is becoming as important as the base model. <a href="https://x.com/dair_ai/status/2075241322655727682">@dair_ai</a> highlighted a study where changing only the orchestration layer cut <strong>blended cost per task 41%</strong>, <strong>tokens 38%</strong>, and <strong>median wall-clock 44%</strong> at quality parity</p></li><li><p>LangChain/LangSmith tooling updates focused on observability for coding agents: tracing <strong>Claude Code</strong> sessions into LangSmith via <a href="https://x.com/LangChain/status/2075233516380717246">@LangChain</a>, plus discussion of <strong>OpenWiki Brains</strong> for proactive memory agents from <a href="https://x.com/BraceSproul/status/2075277759937695979">@BraceSproul</a>, <a href="https://x.com/hwchase17/status/2075277641066938454">@hwchase17</a>, and <a href="https://x.com/colifran_/status/2075406926087934376">@colifran_</a></p></li><li><p><a href="https://x.com/ManusAI/status/2075236343429599432">@ManusAI</a> launched <strong>Branch</strong>, allowing parallel sessions that inherit full context</p></li><li><p><a href="https://x.com/antigravity/status/2075265852992057448">@antigravity</a> described investment in <strong>dynamic agent teams, active sidecars, and generative UI</strong></p></li><li><p><a href="https://x.com/CoreWeave/status/2075293731998286263">@CoreWeave</a> introduced <strong>ARIA</strong>, an AI Research and Improvement Agent inside W&amp;B that reads runs, forms hypotheses, launches experiments, and scores against baselines</p></li><li><p><a href="https://x.com/TheTuringPost/status/2075303983422578740">@TheTuringPost</a> highlighted <strong>SkillCenter</strong>, a package manager/index for agent skills, while <a href="https://x.com/steveruizok/status/2075303919664734295">@steveruizok</a> shipped a &#8220;papercuts&#8221; CLI for agents to report broken tool paths and frustrations</p></li></ul><p><strong>Inference, efficiency, and open model infrastructure</strong></p><ul><li><p><strong>Ollama</strong> announced fundraising and said it now has <strong>9M+ active builders</strong>, framing the moment as scaling &#8220;open models into AI that you can own,&#8221; via <a href="https://x.com/ollama/status/2075211168407503016">@ollama</a></p></li><li><p><strong>Hugging Face / Reachy Mini</strong> economics were striking: <a href="https://x.com/andimarafioti/status/2075222463777042454">@andimarafioti</a> said <strong>9k Reachy Minis</strong> generate <strong>15k hours of conversation/month</strong>; using GPT-realtime would cost <strong>$45k/month</strong>, so they built an open alternative at <strong>$0.25/hour</strong> and free on laptop</p></li><li><p><a href="https://x.com/dmitrshvets/status/2075248269580538081">@dmitrshvets</a> shared speculative decoding research claiming <strong>4.37&#215;</strong> speedup over autoregressive decoding and <strong>+24.7%</strong> over a strong DFlash baseline</p></li><li><p><a href="https://x.com/fal/status/2075284936756539813">@fal</a> detailed a diffusion serving stack reaching <strong>0.45s inference</strong> using kernel optimizations, quantization-aware distillation, and timestep distillation</p></li><li><p><a href="https://x.com/ostrisai/status/2075286667456582080">@ostrisai</a> added isolated reference-token attention for Krea2 edit training; example timings showed major gains from KV caching, such as <strong>31.63s &#8594; 10.90s</strong> for 3 refs</p></li><li><p><a href="https://x.com/vllm_project/status/2075301430123176037">@vllm_project</a> announced the first <strong>vLLM Conference</strong>, underscoring how open inference stacks remain a central layer of the ecosystem</p></li><li><p><a href="https://x.com/QuixiAI/status/2075418782470643958">@QuixiAI</a> reported <strong>Qwen3.6-35B-A3B-NVFP4</strong> at <strong>65 tok/s</strong> on dual B60 with custom SYCL kernels and <strong>128k context</strong></p></li></ul><p><strong>Robotics, multimodal systems, and AI-for-science</strong></p><ul><li><p><a href="https://x.com/perceptroninc/status/2075261142038196727">@perceptroninc</a> launched <strong>Perceptron Egocentric</strong>, an embodied reasoning/annotation system said to beat pipelines built on <strong>Gemini 3.5 Flash</strong> and <strong>Gemini Robotics-ER 1.6</strong></p></li><li><p><a href="https://x.com/DataChaz/status/2075303718153789944">@DataChaz</a> summarized the economics: <strong>10&#8211;15&#215; cheaper</strong> than human annotation, with <strong>+77% end-to-end F1</strong> on <strong>WGO-Bench</strong> (<strong>0.280 vs 0.158</strong>)</p></li><li><p><a href="https://x.com/rohanpaul_ai/status/2075286203583398181">@rohanpaul_ai</a> emphasized the output structure: subtask boundaries, per-hand actions, left/right hand grounding, and dense labels from raw egocentric/robot video</p></li><li><p>Google Research released <strong>SensorFM</strong>, a sensor foundation model trained on <strong>1 trillion minutes</strong> of unlabeled wearable data from <strong>5 million consented participants</strong>, via <a href="https://x.com/GoogleResearch/status/2075283854093607016">@GoogleResearch</a></p></li><li><p><a href="https://x.com/SebastienBubeck/status/2075407986772861047">@SebastienBubeck</a> said GPT&#8209;5.6 helped formalize the <strong>unit distance solution</strong> in <strong>1 million lines of LEAN</strong>, compressing what would previously require a team over years into a short single-person effort</p></li><li><p><a href="https://x.com/TheTuringPost/status/2075289747875107013">@TheTuringPost</a> highlighted a Stanford paper on the <strong>&#8220;Agentic Garden of Forking Paths&#8221;</strong>, where AI research personas reproduced human-like ideological variation; <strong>86%</strong> of analyses passed independent AI review and <strong>78%</strong> were judged methodologically sound by humans</p></li></ul><p><strong>Policy, safety, and ecosystem debate</strong></p><ul><li><p>A cluster of posts sharply criticized the EU&#8217;s <strong>Chat Control</strong> law/proposal from civil-liberties and anti-surveillance angles, including <a href="https://x.com/perrymetzger/status/2075226601298514418">@perrymetzger</a>, <a href="https://x.com/IterIntellectus/status/2075258469561844112">@IterIntellectus</a>, and <a href="https://x.com/dhh/status/2075295777673634256">@dhh</a></p></li><li><p>Open-source advocacy remained loud: <a href="https://x.com/AndrewYNg/status/2075271586400403567">@AndrewYNg</a> said protecting open source AI is critical to permissionless innovation, while <a href="https://x.com/Dan_Jeffries1/status/2075253735563886595">@Dan_Jeffries1</a> argued restricting open source AI would be &#8220;civilizational suicide&#8221;</p></li><li><p><a href="https://x.com/cognition/status/2075308920755618144">@cognition</a> addressed trustworthiness concerns around open-source-derived coding agents, saying their <strong>SWE&#8209;1.7</strong> built on <strong>Kimi K2.7</strong> was specifically trained for trustworthiness and refused surveillance-style scenarios where the base model complied</p></li><li><p>On evaluation methodology and behavior science, <a href="https://x.com/TransluceAI/status/2075271925665063046">@TransluceAI</a> argued for measuring <strong>how systems behave in the world</strong>, not just raw capabilities</p></li><li><p>Forecasting/futures discussion centered on <strong>AI 2040</strong>, with endorsements and critiques from <a href="https://x.com/NeelNanda5/status/2075271483207872874">@NeelNanda5</a>, <a href="https://x.com/RichardMCNgo/status/2075301126921175166">@RichardMCNgo</a>, <a href="https://x.com/scaling01/status/2075296890325712944">@scaling01</a>, and others debating compute gaps, geopolitical assumptions, and takeoff dynamics</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. Chinese Open Models: Releases and Scrutiny</strong></h3><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-openai-launches-gpt-56-solterraluna">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition]]></title><description><![CDATA[SpaceXAI continues to move faster than any other frontier lab on earth.]]></description><link>https://www.latent.space/p/ainews-spacexai-launches-grok-45</link><guid isPermaLink="false">https://www.latent.space/p/ainews-spacexai-launches-grok-45</guid><pubDate>Thu, 09 Jul 2026 06:05:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8D6O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMuQw2BXUAAJaQd.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As <a href="https://x.com/openai/status/2074704958419792299">GPT 5.6 is confirmed to launch tomorrow</a>, today is pretty much the last day anyone will be excited about a GPT 5.5 equivalent model launch, and that is exactly what SpaceXAI did:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;,&quot;full_text&quot;:&quot;We've partnered with SpaceXAI to train Grok 4.5.\n\nIt&#8217;s our most powerful model yet and the first we've built for more than software engineering. &quot;,&quot;username&quot;:&quot;cursor_ai&quot;,&quot;name&quot;:&quot;Cursor&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1970182748146180096/dhZeXi_X_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-08T17:57:18.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HMuQw2BXUAAJaQd.png&quot;,&quot;link_url&quot;:&quot;https://t.co/U4B8Tedl34&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:605,&quot;retweet_count&quot;:1258,&quot;like_count&quot;:14416,&quot;impression_count&quot;:3237140,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>The new <a href="https://cursor.com/blog/grok-4-5">Grok 4.5</a> is a <a href="https://x.com/cursor_ai/status/2074915748544217188">different weight class</a> than the Composer series (<a href="https://x.com/ArtificialAnlys/status/2074956932289282087">1.5T</a>) and despite the solid evals still performs very comparably to the current workhorse Opus and GPTs, although per OpenAI&#8217;s evals team even <a href="https://news.ycombinator.com/item?id=48837396">the mighty SWE-Bench Pro is now saturated/terminally flawed</a> - leaving presumably a small list of successors including <a href="https://www.latent.space/p/ainews-frontiercode-benchmarking">FrontierCode</a>.</p><p>As for training and data disclosures, this is all the information we have.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!beuF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!beuF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 424w, https://substackcdn.com/image/fetch/$s_!beuF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 848w, https://substackcdn.com/image/fetch/$s_!beuF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 1272w, https://substackcdn.com/image/fetch/$s_!beuF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!beuF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png" width="1456" height="1544" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1544,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:461254,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/206247062?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!beuF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 424w, https://substackcdn.com/image/fetch/$s_!beuF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 848w, https://substackcdn.com/image/fetch/$s_!beuF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 1272w, https://substackcdn.com/image/fetch/$s_!beuF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dad0acc-e1ba-4ce4-9196-4ea6e56633a0_1628x1726.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p></p><p></p><blockquote><p>AI News for 7/07/2026-7/08/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Top Story: Grok 4.5 release</strong></p><h2><strong>What happened</strong></h2><p><strong>xAI/&#8220;SpaceXAI&#8221; publicly launched Grok 4.5 as a new coding-and-agents-focused frontier model, positioned on capability-per-dollar rather than absolute benchmark supremacy.</strong></p><ul><li><p>Elon Musk first said Grok 4.5 would be made public &#8220;tomorrow&#8221; based on strong beta feedback, calling it &#8220;Opus-class,&#8221; but faster, more token-efficient, and lower cost <a href="https://x.com/elonmusk/status/2074740539874775163">@elonmusk</a>.</p></li><li><p>Musk later framed Grok 4.5 internally as &#8220;roughly comparable to Opus 4.7, but much faster,&#8221; emphasizing usefulness to Tesla and SpaceX engineers over benchmark chasing <a href="https://x.com/elonmusk/status/2074911038286295049">@elonmusk</a>.</p></li><li><p>The official launch came from xAI&#8217;s account, describing Grok 4.5 as &#8220;our first model trained specifically for coding and agents,&#8221; trained with Cursor, and offering &#8220;frontier intelligence at leading speeds and cost efficiency&#8221; <a href="https://x.com/SpaceXAI/status/2074915721684086811">@SpaceXAI</a>.</p></li><li><p>Cursor said it partnered with xAI to train Grok 4.5, called it &#8220;our most powerful model yet,&#8221; and stressed that it was &#8220;the first we&#8217;ve built for more than software engineering&#8221; <a href="https://x.com/cursor_ai/status/2074915744999969059">@cursor_ai</a>.</p></li><li><p>Cursor also announced in-product availability with &#8220;double usage for the first week&#8221; <a href="https://x.com/cursor_ai/status/2074915747302690991">@cursor_ai</a>.</p></li><li><p>Cursor clarified that &#8220;Grok 4.5 and Composer are two different model weight classes,&#8221; and that Composer 2.5 would remain available with future models in that smaller class <a href="https://x.com/cursor_ai/status/2074915748544217188">@cursor_ai</a>.</p></li><li><p>Early ecosystem support appeared immediately: Grok 4.5 became available in Grok Build/API/Cursor <a href="https://x.com/milichab/status/2074916029848027636">@milichab</a>, day-0 support was announced for Hermes Agent <a href="https://x.com/Teknium/status/2074823590365860254">@Teknium</a>, and later live availability in Hermes Agent/Portal/OpenRouter/Grok subscriptions was confirmed <a href="https://x.com/Teknium/status/2074943072761471314">@Teknium</a>.</p></li><li><p>Musk said the context window would likely move from 500k back to 1M &#8220;by next week&#8221; <a href="https://x.com/elonmusk/status/2074963933199282491">@elonmusk</a>.</p></li></ul><h2><strong>Official claims and product details</strong></h2><h3><strong>Positioning</strong></h3><p>Officially, xAI&#8217;s message was not &#8220;best overall model,&#8221; but near-Opus quality with materially better economics and speed:</p><ul><li><p>&#8220;Opus-class model, but faster, more token-efficient and lower cost&#8221; <a href="https://x.com/elonmusk/status/2074740539874775163">@elonmusk</a></p></li><li><p>&#8220;First model trained specifically for coding and agents&#8221; <a href="https://x.com/SpaceXAI/status/2074915721684086811">@SpaceXAI</a></p></li><li><p>&#8220;Frontier intelligence at leading speeds and cost efficiency&#8221; <a href="https://x.com/SpaceXAI/status/2074915721684086811">@SpaceXAI</a></p></li><li><p>&#8220;Most powerful model yet&#8221; and &#8220;first we&#8217;ve built for more than software engineering&#8221; <a href="https://x.com/cursor_ai/status/2074915744999969059">@cursor_ai</a></p></li></ul><p>This framing matters: xAI is explicitly targeting the coding-agent workflow market that has recently been dominated by Anthropic/OpenAI/Cursor-style tool-using systems, not just general chat.</p><h3><strong>Pricing and context</strong></h3><p>The concrete numbers that surfaced:</p><ul><li><p>Official pricing: <strong>$2 / 1M input tokens, $6 / 1M output tokens</strong> <a href="https://x.com/scaling01/status/2074914032880947601">@scaling01</a></p></li><li><p>Artificial Analysis repeated the same price point and added:</p><ul><li><p><strong>cache hits discounted by 75% to $0.5 / 1M tokens</strong></p></li><li><p><strong>long inputs over 200k tokens cost double</strong></p></li><li><p><strong>500k context window</strong>, down from Grok 4.3&#8217;s <strong>1M</strong></p></li><li><p><strong>vision input retained</strong></p></li><li><p><strong>configurable reasoning retained</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li></ul></li><li><p>Musk later said the context window would probably upgrade back to <strong>1M</strong> soon <a href="https://x.com/elonmusk/status/2074963933199282491">@elonmusk</a>.</p></li></ul><p>Relative pricing comparisons cited by users:</p><ul><li><p>Grok 4.5: <strong>$2 in / $6 out</strong></p></li><li><p>GPT-5.6: <strong>$5 in / $30 out</strong></p></li><li><p>Opus 4.8: <strong>$5 in / $25 out</strong> <a href="https://x.com/kimmonismus/status/2074940669718638780">@kimmonismus</a></p></li></ul><h3><strong>Model size</strong></h3><p>One important spec surfaced via third-party reporting of Musk&#8217;s disclosure:</p><ul><li><p>Grok 4.5 is <strong>3x larger than Grok 4.3 at 1.5T parameters</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li></ul><p>That is a notable jump, and likely central to why multiple observers interpreted 4.5 as xAI&#8217;s first entry into the true flagship coding-agent tier rather than an iterative refresh.</p><h2><strong>Benchmarks and independent evaluations</strong></h2><h3><strong>Artificial Analysis</strong></h3><p>Artificial Analysis provided the most substantive external evaluation in the tweet set.</p><p>Key results:</p><ul><li><p><strong>#4 on Artificial Analysis Intelligence Index</strong>, score <strong>54</strong>, behind only <strong>Fable 5, GPT-5.5, and Opus 4.8</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>+16 points vs Grok 4.3</strong> on the same index <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>GDPval-AA v2 Elo 1543</strong>, also ranking <strong>#4</strong>, behind Anthropic&#8217;s latest Claude releases <a href="https://x.com/ArtificialAnlys/status/2074942097158021371">@ArtificialAnlys</a></p></li><li><p><strong>Top score on &#964;&#179;-Banking: 33%</strong>, above <strong>31% for GPT-5.5 (xhigh)</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>Artificial Analysis Coding Agent Index score 76</strong> in Grok Build, &#8220;on par with GPT-5.5 in Codex&#8221; and below Fable 5 in Claude Code <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>Cost per Intelligence Index task: $0.31</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>Cost per GDPval task: $0.49</strong> <a href="https://x.com/ArtificialAnlys/status/2074942097158021371">@ArtificialAnlys</a></p></li><li><p><strong>Cost per Coding Agent Index task: $2.59</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>Average output tokens per Intelligence Index task: ~14k</strong>, over <strong>60% lower than Opus 4.8</strong> <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li><li><p><strong>Average total tokens per Coding Agent Index task: 1.9M</strong>, versus <strong>7.2M</strong> for Fable 5 in Claude Code and <strong>6.2M</strong> for GPT-5.5 in Codex <a href="https://x.com/ArtificialAnlys/status/2074956932289282087">@ArtificialAnlys</a></p></li></ul><p>Artificial Analysis&#8217; interpretation was clear: Grok 4.5 is near-frontier on capability, but unusually strong on efficiency, making it sit on the Pareto frontier for cost/performance.</p><p>Musk explicitly amplified the Artificial Analysis assessment <a href="https://x.com/elonmusk/status/2074948489792860456">@elonmusk</a>.</p><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-spacexai-launches-grok-45">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO]]></title><description><![CDATA[2 years after our first coverage, we return with Modal's other cofounder to explore why Agent Experience is working now, and everything they have learned building the new agent cloud.]]></description><link>https://www.latent.space/p/modal2026</link><guid isPermaLink="false">https://www.latent.space/p/modal2026</guid><pubDate>Wed, 08 Jul 2026 22:55:07 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/205716015/df6332898c08341f79886bb62af7a458.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We&#8217;ve been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from <a href="https://www.latent.space/p/databricks">Databricks</a> to <a href="https://www.latent.space/p/daytona">Daytona</a> to <a href="https://www.latent.space/p/railway">Railway</a> and, even further back, <a href="https://www.latent.space/p/e2b?utm_source=publication-search">E2B</a>, but we&#8217;re excited to conclude this series returning to Modal, which has just raised a monster <a href="https://modal.com/blog/modal-series-c">$355M Series C</a>.</p><p>The cloud was built for developers. But <strong>agents are now changing that.</strong></p><p>The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.</p><p><strong>However, agents don&#8217;t have that luxury. </strong>Now in this new era of agents, everything has to be tighter.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a5RX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a5RX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 424w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 848w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 1272w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a5RX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png" width="1456" height="224" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9828917e-57a7-443a-983c-da258c66a614_1714x264.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:224,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87253,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/205716015?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a5RX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 424w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 848w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 1272w, https://substackcdn.com/image/fetch/$s_!a5RX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9828917e-57a7-443a-983c-da258c66a614_1714x264.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption"><a href="https://modal.com/blog/agents-devex">Agents need good developer experience too - Modal Blog</a></figcaption></figure></div><p>They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. <strong>Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly.</strong> Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/akshat_b/status/2044163135624347882&quot;,&quot;full_text&quot;:&quot;Pretty cool watching agents fully realize the dream of programmatic infra. Give your agent the gift of better primitives today, and see how much more they can get done!&quot;,&quot;username&quot;:&quot;akshat_b&quot;,&quot;name&quot;:&quot;Akshat Bubna&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1971998467359735808/QDSXHs4W_normal.jpg&quot;,&quot;date&quot;:&quot;2026-04-14T21:17:24.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;&quot;,&quot;username&quot;:&quot;modal&quot;,&quot;name&quot;:&quot;Modal&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1899521344766627840/3lInURMw_normal.jpg&quot;},&quot;reply_count&quot;:3,&quot;retweet_count&quot;:6,&quot;like_count&quot;:61,&quot;impression_count&quot;:7624,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;33bea364-00cc-411a-939b-617e4b343a63&quot;,&quot;caption&quot;:&quot;We&#8217;re writing this one day after the monster release of OpenAI&#8217;s Sora and Gemini 1.5. We covered this on Alex Volkov &#8216;s ThursdAI space, so head over there for our takes.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal&quot;,&quot;publishedBylines&quot;:[],&quot;post_date&quot;:&quot;2024-02-16T17:42:47.047Z&quot;,&quot;cover_image&quot;:&quot;https://substack-video.s3.amazonaws.com/video_upload/post/141683117/d7f5c748-0f58-453e-b1bb-16eea4c46b7d/transcoded-1708369328.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.latent.space/p/modal&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:141683117,&quot;type&quot;:&quot;podcast&quot;,&quot;reaction_count&quot;:18,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1084089,&quot;publication_name&quot;:&quot;Latent.Space&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!DbYa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b0838a-bd14-46a1-801c-b6a2046e5c1e_1130x1130.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>At the time, Modal was just a teeny little company with a <a href="https://tracxn.com/d/companies/modal/__tHK2ShUcB0Q1o6j-hbJ-xcZMxDsw0P3kCJ85veVeYjU">$17M Series A</a>.</p><p>Today, fresh off their <strong>$355M Series C</strong>, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as <a href="https://modal.com/products/inference">elastic inference</a>, <a href="https://modal.com/products/sandboxes">sandboxes</a>, GPU burst, post-training, background agents, and <a href="https://modal.com/solutions/coding-agents">infrastructure that agents themselves can operate</a>.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/akshat_b/status/2057546075888586795?s=20&quot;,&quot;full_text&quot;:&quot;Raising $ is cool. What&#8217;s even cooler is getting to work every day with this incredible group of humans.\n\nWe like solving hard problems and building things we can be proud of. If this is you, come join us! We&#8217;re just getting started :)&quot;,&quot;username&quot;:&quot;akshat_b&quot;,&quot;name&quot;:&quot;Akshat Bubna&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1971998467359735808/QDSXHs4W_normal.jpg&quot;,&quot;date&quot;:&quot;2026-05-21T19:36:26.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HI3azRXW0AAbBYJ.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/DPe0fcKHvs&quot;}],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;&quot;,&quot;username&quot;:&quot;modal&quot;,&quot;name&quot;:&quot;Modal&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1899521344766627840/3lInURMw_normal.jpg&quot;},&quot;reply_count&quot;:19,&quot;retweet_count&quot;:20,&quot;like_count&quot;:283,&quot;impression_count&quot;:51718,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>In this episode, <strong>Modal CTO Akshat Bubna</strong> joins swyx and Vibhu to unpack why AI applications don&#8217;t fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from <strong>developer experience to agent experience</strong>.</p><p>We go deep on Modal&#8217;s AI infra stack: serverless functions, decorator-based infrastructure, <strong>elastic inference for custom models</strong>, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal&#8217;s capacity pool across <strong>17 cloud providers</strong>. Akshat also explains why RL rollouts can require <strong>100,000 sandboxes</strong>, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.</p><div><hr></div><h3>We discuss:</h3><ul><li><p>Why <strong>Kubernetes</strong> wasn&#8217;t built for <strong>bursty AI workloads</strong></p></li><li><p>How Modal started as a <strong>better runtime</strong> before becoming an <strong>AI cloud</strong></p></li><li><p>Why Modal added <strong>GPUs before ChatGPT</strong></p></li><li><p>The shift from <strong>developer experience</strong> to <strong>agent experience</strong></p></li><li><p>Why <strong>observability</strong> matters when agents are writing the code</p></li><li><p><strong>Elastic inference</strong> for custom models across audio, video, robotics, and comp bio</p></li><li><p><strong>GPU snapshotting</strong>, cold starts, and why inference workloads are so bursty</p></li><li><p>Why <strong>RL rollouts</strong> can require <strong>100,000 sandboxes</strong></p></li><li><p><strong>DeFlash</strong>, speculative decoding, and frontier-level inference performance</p></li><li><p><strong>Auto Endpoints</strong> and making optimized inference easier to deploy</p></li><li><p>What Modal adds beyond <strong>vLLM</strong>, <strong>SGLang</strong>, and raw GPU rental</p></li><li><p>Modal&#8217;s <strong>17-cloud</strong> capacity pool and <strong>supercloud</strong> strategy</p></li><li><p><strong>Networked sandboxes</strong>, sidecars, private IPv6, and RDMA</p></li><li><p><strong>Serverless multi-node training</strong> for post-training and research workloads</p></li><li><p><strong>Auto-research</strong>, model-guided sweeps, and agents launching GPU experiments</p></li><li><p><strong>Compute strategy</strong>, capacity planning, and batch tiers</p></li><li><p>Why production agents need <strong>specialized sandboxes</strong> and <strong>hard guardrails</strong></p></li><li><p>Modal&#8217;s take on <strong>managed agents</strong>, <strong>CI</strong>, Gitpod/Ona, Python, TypeScript, and Modal Bench</p></li></ul><div><hr></div><p><strong>Akshat Bubna</strong></p><ul><li><p><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/akshat-bubna-188885103">https://www.linkedin.com/in/akshat-bubna-188885103</a></p></li><li><p><strong>X:</strong> <a href="https://x.com/akshat_b">https://x.com/akshat_b</a></p></li></ul><p><strong>Modal</strong></p><ul><li><p><strong>Website:</strong> <a href="https://modal.com">https://modal.com</a></p></li></ul><div><hr></div><h2>Timestamps</h2><p><strong>00:00:00</strong> Introduction</p><p><strong>00:00:39</strong> Modal&#8217;s origin and why Kubernetes wasn&#8217;t enough</p><p><strong>00:04:32</strong> Developer Experience &#8594; Agent Experience</p><p><strong>00:06:21</strong> Modal&#8217;s AI cloud primitives</p><p><strong>00:09:14</strong> Sandboxes, agent loops, and proto-Cognition</p><p><strong>00:12:12</strong> Elastic inference, GPU snapshotting, and 100,000 sandboxes</p><p><strong>00:15:24</strong> DeFlash, speculative decoding, and Auto Endpoints</p><p><strong>00:19:59</strong> Production-grade inference beyond raw GPUs</p><p><strong>00:22:00</strong> Background agents, Ramp Inspect, and the agent lifecycle</p><p><strong>00:24:08</strong> Modal&#8217;s 17-cloud supercloud strategy</p><p><strong>00:26:40</strong> Networked sandboxes, private IPv6, and RDMA</p><p><strong>00:32:48</strong> Multi-node training, post-training, and auto research</p><p><strong>00:37:36</strong> Compute strategy, capacity planning, and batch tiers</p><p><strong>00:40:55</strong> Open models, real-time AI, and production agent infra</p><p><strong>00:43:06</strong> Hard guardrails, managed agents, and specialized sandboxes</p><p><strong>00:46:06</strong> Why AI made infrastructure exciting again</p><p><strong>00:48:30</strong> Model APIs, differentiated products, and agentic video</p><p><strong>00:51:50</strong> CI, coding-agent infra, SDKs, and Modal Bench</p><p><strong>00:57:28</strong> Closing Thoughts</p><div><hr></div><h1>Transcript</h1><h2>Introduction: Modal, Series C, and the Art Party</h2><p><strong>Swyx [00:00:00]:</strong> We&#8217;re here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.</p><p><strong>Akshat [00:00:10]:</strong> Thank you.</p><p><strong>Swyx [00:00:11]:</strong> Your party yesterday was amazing.</p><p><strong>Akshat [00:00:15]:</strong> Yeah.</p><p><strong>Swyx [00:00:15]:</strong> From all the photos and all the swag.</p><p><strong>Akshat [00:00:17]:</strong> We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.</p><p><strong>Swyx [00:00:25]:</strong> Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.</p><h2>Modal&#8217;s Origin: A New Runtime Beyond Kubernetes</h2><p><strong>Akshat [00:00:39]:</strong> I first met Eric, who&#8217;s the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It&#8217;s because you have to run them on Kubernetes. Kubernetes is hard to manage. It&#8217;s not built for burstiness and, custom images,</p><p><strong>Swyx [00:01:03]:</strong> Yeah</p><p><strong>Akshat [00:01:03]:</strong> It has a terrible developer experience.</p><p><strong>Swyx [00:01:05]:</strong> And I&#8217;ll, I&#8217;ll interject</p><p><strong>Akshat [00:01:06]:</strong> Yeah</p><p><strong>Swyx [00:01:07]:</strong> For listeners, who are new, we interviewed Eric two years ago, and there&#8217;s a bit more of the story there from Spotify and all those things.</p><p><strong>Swyx [00:01:14]:</strong> And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, &#8220;Okay, I need to take Modal very seriously&#8221; moment.</p><p><strong>Akshat [00:01:26]:</strong> Yeah.</p><p><strong>Swyx [00:01:26]:</strong> But it was still very unclear, like, do I need all this for just my data pipelines?</p><p><strong>Akshat [00:01:33]:</strong> Yeah. initially what we were thinking about was if we build a better runtime, it&#8217;s a very useful primitive in itself. It&#8217;s There&#8217;s a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would&#8217;ve been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.</p><h2>From Serverless Containers to GPU Workloads</h2><p><strong>Swyx [00:02:19]:</strong> Nice.</p><p><strong>Akshat [00:02:19]:</strong> We just didn&#8217;t think it would be that big of a deal.</p><p><strong>Swyx [00:02:22]:</strong> Yeah, just like add A100.</p><p><strong>Vibhu [00:02:23]:</strong> Was there any, like, early key problem that really sparked off why you built it?</p><p><strong>Akshat [00:02:28]:</strong> Yeah. Primarily it&#8217;s just, none of the tooling that was out there was built for, one, a really great developer experience, and also there&#8217;s a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there&#8217;s just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.</p><h2>Software-Defined Infrastructure and Decorator-Based DX</h2><p><strong>Swyx [00:03:13]:</strong> Yeah. Yeah. Be nice. I don&#8217;t know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.</p><p><strong>Akshat [00:03:22]:</strong> Yeah, the self-provisioning</p><p><strong>Swyx [00:03:23]:</strong> Self-provisioning.</p><p><strong>Akshat [00:03:24]:</strong> Yeah.</p><p><strong>Swyx [00:03:24]:</strong> Yeah. I can&#8217;t even remember my own post.</p><p><strong>Swyx [00:03:26]:</strong> And then you put me on the landing page.</p><p><strong>Akshat [00:03:28]:</strong> Yeah. We really like, the term and so we stole it.</p><p><strong>Swyx [00:03:32]:</strong> Because you had the insight that everything can just be in decorators co-located with the code, right?</p><p><strong>Akshat [00:03:37]:</strong> Yeah.</p><p><strong>Swyx [00:03:37]:</strong> Was that a big part of the original</p><p><strong>Akshat [00:03:39]:</strong> Yes</p><p><strong>Swyx [00:03:39]:</strong> Story or it was just like a DX layer?</p><p><strong>Akshat [00:03:41]:</strong> That was, really important because we really didn&#8217;t want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you&#8217;re doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that&#8217;s more expressive and dynamic. and so yeah, that was always a very important part.</p><p><strong>Swyx [00:04:04]:</strong> Then the pushback is this is a DSL.</p><p><strong>Akshat [00:04:07]:</strong> Yeah.</p><p><strong>Swyx [00:04:07]:</strong> It&#8217;s you&#8217;re closed source. I am locked into Modal.</p><p><strong>Akshat [00:04:11]:</strong> Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you&#8217;re using, how you&#8217;re scaling things up, but you still own the code.</p><p><strong>Akshat [00:04:27]:</strong> And that&#8217;s, that&#8217;s been an important, part of our story, even as we do inference now.</p><p><strong>Swyx [00:04:32]:</strong> Yeah.</p><p><strong>Vibhu [00:04:32]:</strong> How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there&#8217;s very agent native primitives that are different than if I&#8217;m doing this myself, right?</p><h2>Developer Experience &#8594; Agent Experience</h2><p><strong>Akshat [00:04:54]:</strong> We&#8217;ve changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that&#8217;s not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.</p><p><strong>Swyx [00:05:34]:</strong> Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.</p><p><strong>Akshat [00:05:38]:</strong> Yeah.</p><p><strong>Swyx [00:05:38]:</strong> Well, the negative thesis now is that nobody&#8217;s looking at their code anymore, so there&#8217;s no point.</p><p><strong>Akshat [00:05:44]:</strong> Yeah, people aren&#8217;t looking at code. one thing we still see is really important is observability.</p><p><strong>Swyx [00:05:51]:</strong> Yeah.</p><p><strong>Akshat [00:05:51]:</strong> Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what&#8217;s going on and, make judgment calls and whatnot. and that&#8217;s I feel like, Maybe more important now than looking at the code itself.</p><p><strong>Swyx [00:06:11]:</strong> Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.</p><h2>What Modal Is For: AI Cloud Primitives</h2><p><strong>Akshat [00:06:21]:</strong> Yeah.</p><p><strong>Swyx [00:06:22]:</strong> So I think it takes a bit of restraint to not specialize, to say, &#8220;I want to ship a new primitive,&#8221; and then just be general purpose.</p><p><strong>Swyx [00:06:31]:</strong> People ask you, &#8220;What are you for?&#8221; You&#8217;re like, &#8220; I don&#8217;t know. We can do this, we can do that.&#8221;</p><p><strong>Vibhu [00:06:36]:</strong> Well, I&#8217;d be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There&#8217;s a lot you guys do, sandboxes, GPUs, everything. How do you answer?</p><p><strong>Akshat [00:06:46]:</strong> Modal is a cloud platform that&#8217;s built for, where we&#8217;ve built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.</p><p><strong>Akshat [00:07:00]:</strong> But we&#8217;re building a lot more</p><p><strong>Swyx [00:07:02]:</strong> I noticed you didn&#8217;t say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.</p><p><strong>Akshat [00:07:09]:</strong> Yeah, absolutely. We&#8217;re, we&#8217;re not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they&#8217;re, they&#8217;re, they&#8217;re just shaped differently.</p><h2>Working Alongside Frontier Startups</h2><p><strong>Vibhu [00:07:26]:</strong> I think you&#8217;re building a lot of it alongside the startups, right? They&#8217;re innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they&#8217;re, they&#8217;re innovating with you, right? And that&#8217;s not something AWS is doing directly with.</p><p><strong>Akshat [00:07:45]:</strong> Yeah, absolutely. I think, this is again classic. We&#8217;re a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.</p><p><strong>Swyx [00:07:54]:</strong> So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, &#8220;What are you doing here?&#8221; They&#8217;re like, &#8220;Yeah, I just. I am embedded inside of Cog.&#8221;</p><p><strong>Akshat [00:08:05]:</strong> Yeah, I think that was Peyton. We sent him over</p><p><strong>Swyx [00:08:07]:</strong> Yeah.</p><p><strong>Akshat [00:08:07]:</strong> Because, the latency of communication was too high otherwise.</p><p><strong>Swyx [00:08:12]:</strong> Yeah, distributed node, you have to - you have to place one and collocate.</p><p><strong>Vibhu [00:08:16]:</strong> Yeah.</p><p><strong>Swyx [00:08:16]:</strong> So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, &#8220;Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.&#8221; And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.</p><h2>smol developer, Sandboxes, and Proto-Cognition</h2><p><strong>Akshat [00:08:39]:</strong> Yeah, you blew up on Hacker News and,</p><p><strong>Swyx [00:08:41]:</strong> Yeah</p><p><strong>Akshat [00:08:41]:</strong> We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.</p><p><strong>Swyx [00:08:53]:</strong> Yeah. That - So to me, that was proto-cognition.</p><p><strong>Akshat [00:08:55]:</strong> Right.</p><p><strong>Swyx [00:08:56]:</strong> If only I had, like, stuck to it.</p><p><strong>Swyx [00:08:58]:</strong> Like, that was like, if - did you say draw the tech tree</p><p><strong>Akshat [00:09:00]:</strong> Absolutely</p><p><strong>Swyx [00:09:00]:</strong> You&#8217;re just like, &#8220;Yeah, like, probably this will happen.&#8221;</p><p><strong>Akshat [00:09:02]:</strong> Yeah. Like, he was so close. You were just rebuilding upon us</p><p><strong>Swyx [00:09:04]:</strong> I just didn&#8217;t realize.</p><p><strong>Akshat [00:09:05]:</strong> But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.</p><p><strong>Swyx [00:09:14]:</strong> Yeah.</p><p><strong>Akshat [00:09:14]:</strong> This is like twenty-three.</p><p><strong>Swyx [00:09:15]:</strong> Yeah.</p><p><strong>Akshat [00:09:16]:</strong> So we built</p><p><strong>Swyx [00:09:17]:</strong> You introduced a new API right after that.</p><p><strong>Akshat [00:09:18]:</strong> Yeah.</p><p><strong>Swyx [00:09:19]:</strong> Yes.</p><p><strong>Akshat [00:09:19]:</strong> Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developer</p><p><strong>Swyx [00:09:28]:</strong> Smol developer</p><p><strong>Akshat [00:09:28]:</strong> And put it in a loop, so the agent can iterate on itself.</p><p><strong>Swyx [00:09:33]:</strong> Loops are hot these days.</p><p><strong>Vibhu [00:09:34]:</strong> It&#8217;s the looper.</p><p><strong>Akshat [00:09:34]:</strong> Yeah.</p><p><strong>Vibhu [00:09:35]:</strong> Loops in. When was this, twenty-three?</p><p><strong>Akshat [00:09:38]:</strong> Yeah.</p><p><strong>Vibhu [00:09:39]:</strong> A small check.</p><p><strong>Akshat [00:09:39]:</strong> Yeah.</p><p><strong>Swyx [00:09:39]:</strong> It&#8217;s like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?</p><p><strong>Swyx [00:09:46]:</strong> Like, you&#8217;re just trying to like. They&#8217;re not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.</p><p><strong>Akshat [00:09:55]:</strong> Yeah.</p><p><strong>Akshat [00:09:55]:</strong> I don&#8217;t remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.</p><p><strong>Swyx [00:10:03]:</strong> Yeah. But like, then. So okay, like now talking to myself three years ago, the answer</p><p><strong>Vibhu [00:10:08]:</strong> Of course they will get better</p><p><strong>Swyx [00:10:09]:</strong> Collect all the failures, build benchmark, and then collect all the, examples, build the RL environment</p><p><strong>Akshat [00:10:15]:</strong> Right</p><p><strong>Swyx [00:10:15]:</strong> Sell it for like ten billion dollars to Meta.</p><p><strong>Swyx [00:10:17]:</strong> And then also train a model and then sell that for sixty billion dollars to Elon. And this is</p><p><strong>Akshat [00:10:23]:</strong> Yeah, of course</p><p><strong>Swyx [00:10:23]:</strong> The funny machine. Like, it&#8217;s like, it&#8217;s about the hardware.</p><p><strong>Akshat [00:10:28]:</strong> It&#8217;s hard to have that inherent conviction that the stuff will get that much better.</p><p><strong>Swyx [00:10:33]:</strong> In retrospect, it&#8217;s so fucking obvious.</p><p><strong>Akshat [00:10:36]:</strong> Fair enough.</p><p><strong>Swyx [00:10:37]:</strong> Like, what else were we doing back then? I don&#8217;t know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn&#8217;t blow up until, like, last year.</p><p><strong>Akshat [00:10:49]:</strong> Yeah.</p><p><strong>Swyx [00:10:50]:</strong> So there was like a couple years of quietness.</p><p><strong>Akshat [00:10:52]:</strong> Exactly, yeah. We were</p><p><strong>Vibhu [00:10:53]:</strong> I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he&#8217;s like, &#8220; there&#8217;s this cool company, Modal. They&#8217;ll like spin up a GPU sandbox, we can throw it on there. They&#8217;ll take a Hugging Face link.&#8221; And like there&#8217;s so much value just right there, right? Like instant hosting, spin it up, spin it down. It&#8217;ll stay cold, but we run the demo a few days later, it&#8217;ll come back up and like all this stuff in retrospect, like it&#8217;s still what we needed like today.</p><p><strong>Akshat [00:11:27]:</strong> Yeah, it&#8217;s still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it&#8217;s it&#8217;s not about scaling from zero to one, but it&#8217;s how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It&#8217;s the same shape problem.</p><h2>Elastic Inference, GPU Autoscaling, and Custom Models</h2><p><strong>Vibhu [00:11:50]:</strong> Okay. So you look at, say, Cursor Composer, right?</p><p><strong>Akshat [00:11:53]:</strong> Yeah.</p><p><strong>Vibhu [00:11:53]:</strong> They had a. &#8220;We&#8217;ll do RL on a model every couple hours.&#8221; you guys have a whole version of RL inference gym and whatnot.</p><p><strong>Vibhu [00:12:01]:</strong> When you look at workloads like that, you&#8217;re doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That&#8217;s the example for we do need it, right?</p><p><strong>Akshat [00:12:12]:</strong> Yeah. Well, so I&#8217;ll, I&#8217;ll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it&#8217;s like diurnal. It&#8217;s Some days, like the company will do a launch and, they&#8217;ll need like, way more. And it&#8217;s not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.</p><p><strong>Akshat [00:13:20]:</strong> So that&#8217;s like our sort</p><p><strong>Vibhu [00:13:22]:</strong> And that</p><p><strong>Akshat [00:13:22]:</strong> Yeah</p><p><strong>Vibhu [00:13:22]:</strong> That in and of itself is a huge category. There&#8217;s a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that&#8217;s carved into its own niche for language models, at least right now.</p><p><strong>Akshat [00:13:36]:</strong> Yeah. the thing that we have specialized in is the autoscaling aspect.</p><p><strong>Vibhu [00:13:41]:</strong> Yeah.</p><p><strong>Akshat [00:13:41]:</strong> Because we found that it&#8217;s not universally true that everyone else can autoscale, and we&#8217;ve gone deeper into it on the tech side by, we&#8217;ve incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it&#8217;s That&#8217;s why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we&#8217;ll see, a lot of companies, before they have a training run, they&#8217;ll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you&#8217;re doing RL. RL is just</p><h2>RL, Batch Jobs, and 100,000 Sandboxes</h2><p><strong>Vibhu [00:14:28]:</strong> Or commerce</p><p><strong>Akshat [00:14:28]:</strong> Insanely bursty.</p><p><strong>Vibhu [00:14:29]:</strong> Yeah.</p><p><strong>Akshat [00:14:30]:</strong> Yeah. Like when you&#8217;re doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.</p><p><strong>Vibhu [00:14:37]:</strong> Yeah. I&#8217;m curious if you&#8217;ve seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced this</p><p><strong>Akshat [00:14:45]:</strong> Yeah</p><p><strong>Vibhu [00:14:45]:</strong> They&#8217;re, they&#8217;re trying to do training. That also seems like a different workload, right? If you&#8217;re doing training twenty-four/seven per se, there&#8217;s a very weird dynamic of how you&#8217;re using GPUs between people and whatnot, but seems like something you guys would work for.</p><p><strong>Akshat [00:15:00]:</strong> As you said, we&#8217;re, we&#8217;re fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It&#8217;s possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we&#8217;re, we&#8217;re just waiting to see</p><p><strong>Vibhu [00:15:23]:</strong> Yeah</p><p><strong>Akshat [00:15:23]:</strong> How it shakes out.</p><p><strong>Vibhu [00:15:24]:</strong> Is there a primitive that you added after sandboxing that was the next step in the story?</p><h2>LLM Inference, DeFlash, and Speculative Decoding</h2><p><strong>Akshat [00:15:32]:</strong> I guess we&#8217;ve been going much deeper into LLM inference</p><p><strong>Vibhu [00:15:35]:</strong> Yeah</p><p><strong>Akshat [00:15:35]:</strong> Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren&#8217;t, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we&#8217;ve been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we&#8217;ve open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we&#8217;re thinking about here</p><p><strong>Vibhu [00:16:23]:</strong> I thought this was</p><p><strong>Akshat [00:16:24]:</strong> Yeah</p><p><strong>Vibhu [00:16:24]:</strong> An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.</p><p><strong>Akshat [00:16:33]:</strong> Yeah.</p><p><strong>Vibhu [00:16:33]:</strong> Anything you wanna point out from this around, what people should know?</p><p><strong>Akshat [00:16:39]:</strong> Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.</p><p><strong>Vibhu [00:16:44]:</strong> Yes.</p><p><strong>Akshat [00:16:44]:</strong> I will, yes.</p><p><strong>Vibhu [00:16:45]:</strong> I think, like</p><p><strong>Akshat [00:16:46]:</strong> Yeah</p><p><strong>Vibhu [00:16:46]:</strong> So we&#8217;ve covered like Eagle and all this</p><p><strong>Akshat [00:16:47]:</strong> Yeah</p><p><strong>Vibhu [00:16:47]:</strong> Like Hydra and all those things, but it was like two years ago.</p><p><strong>Akshat [00:16:51]:</strong> Yeah.</p><p><strong>Vibhu [00:16:51]:</strong> I think it doesn&#8217;t hurt, right?</p><p><strong>Akshat [00:16:52]:</strong> Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it&#8217;s faster is if you&#8217;re predicting, one token at once, you&#8217;re bound by memory bandwidth. But if you can batch the verification of, the draft model, then you&#8217;re much more efficient using compute, and it&#8217;s faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that&#8217;s, multiple times of, the original model speed. and well, that&#8217;s what we highlight here. It&#8217;s Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decrease</p><p><strong>Vibhu [00:17:47]:</strong> Like two to four X.</p><p><strong>Akshat [00:17:48]:</strong> Yeah, exactly.</p><p><strong>Vibhu [00:17:48]:</strong> Without much head-on performance.</p><p><strong>Akshat [00:17:50]:</strong> Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,</p><p><strong>Vibhu [00:17:57]:</strong> I meant quality performance</p><p><strong>Akshat [00:17:58]:</strong> Probably not by much</p><p><strong>Vibhu [00:17:58]:</strong> But yeah. I think</p><p><strong>Akshat [00:17:59]:</strong> So there&#8217;s no drop in quality performance</p><p><strong>Vibhu [00:18:01]:</strong> Yeah</p><p><strong>Akshat [00:18:01]:</strong> Because you&#8217;re always. You&#8217;re never accepting a token that the big model</p><p><strong>Vibhu [00:18:04]:</strong> It&#8217;s strictly better</p><p><strong>Akshat [00:18:05]:</strong> Yeah</p><p><strong>Vibhu [00:18:05]:</strong> Or it&#8217;s same.</p><p><strong>Akshat [00:18:06]:</strong> Exactly.</p><p><strong>Vibhu [00:18:07]:</strong> Right. Yeah.</p><p><strong>Akshat [00:18:08]:</strong> And so we&#8217;ve been working a bunch on DeFlash, which is a block-based speculator. so it&#8217;s instead of predicting, one token at a time, it&#8217;s predicting a block. And we&#8217;ve been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it&#8217;s it&#8217;s something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we&#8217;re, we&#8217;re launching is, as you run an auto endpoint, we shadow traffic</p><h2>Auto Endpoints and Frontier-Level Performance</h2><p><strong>Vibhu [00:18:54]:</strong> Do you want to explain what auto endpoints are?</p><p><strong>Akshat [00:18:57]:</strong> Yeah.</p><p><strong>Vibhu [00:18:57]:</strong> I lovely, yeah.</p><p><strong>Akshat [00:18:58]:</strong> Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don&#8217;t wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we&#8217;ve made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there&#8217;s full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It&#8217;s it&#8217;s not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.</p><p><strong>Vibhu [00:19:59]:</strong> I guess just to understand it directly, you have the GPUs, you have an endpoint that&#8217;s compatible, you serve open model. If someone was to do this themselves, what&#8217;s the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What&#8217;s the delta that plugging into something this, like this offers outside of the benefit of, scaling?</p><h2>Production Inference Beyond Raw GPUs</h2><p><strong>Akshat [00:20:34]:</strong> It&#8217;s interesting because we&#8217;ve taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.</p><p><strong>Vibhu [00:21:20]:</strong> Yeah. And I will say it&#8217;s not that straightforward to just. like what I said is easier said than done, right?</p><p><strong>Akshat [00:21:26]:</strong> Yeah.</p><p><strong>Vibhu [00:21:27]:</strong> It&#8217;s I think still for the average person, still hard to just gut check using different. There&#8217;s, there&#8217;s quite a bit of combinations you can make there. the trade-offs aren&#8217;t really known at face value.</p><p><strong>Akshat [00:21:40]:</strong> Yeah. it&#8217;s it&#8217;s not just that. I think it&#8217;s it&#8217;s that running production-grade inference is a hard infer problem.</p><p><strong>Vibhu [00:21:49]:</strong> Yeah</p><p><strong>Akshat [00:21:49]:</strong> Even if you subtract out the autoscaling</p><p><strong>Vibhu [00:21:50]:</strong> Yeah</p><p><strong>Akshat [00:21:51]:</strong> Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.</p><h2>The Model and Agent Lifecycle</h2><p><strong>Vibhu [00:22:00]:</strong> There&#8217;s a lot of innovation that you can do here. I think, it&#8217;s very interesting that you&#8217;re starting to encroach on, like as you become a full cloud, you&#8217;re starting to encroach on other people&#8217;s turf.</p><p><strong>Vibhu [00:22:09]:</strong> What will you not do?</p><p><strong>Akshat [00:22:13]:</strong> Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we&#8217;re focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let&#8217;s say, sandbox, do persistent storage, a whole bunch of other stuff.</p><p><strong>Vibhu [00:22:38]:</strong> We talked to Cole, who did, OpenInspect. Yeah.</p><p><strong>Akshat [00:22:42]:</strong> Yeah.</p><p><strong>Vibhu [00:22:42]:</strong> And RealInspect also is on Modal.</p><p><strong>Akshat [00:22:44]:</strong> Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.</p><h2>Ramp Inspect and Background Agents</h2><p><strong>Vibhu [00:23:02]:</strong> Yeah. That&#8217;s the new CTO of, Ramp right there.</p><p><strong>Akshat [00:23:05]:</strong> Yeah, Rahul.</p><p><strong>Vibhu [00:23:08]:</strong> It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guys</p><h2>The Inference Inflection: CPU, GPU, and Co-Location</h2><p><strong>Vibhu [00:23:19]:</strong> You weren&#8217;t that much in the GPU game, and now you&#8217;re all about, inference. And one of the points that I hinged on for Jensen&#8217;s keynote at GTC this year was, what we&#8217;re calling like the inference inflection, right? That let&#8217;s say in AI workloads or machine learning workloads, it used to be like, let&#8217;s call it eight to one GPU to CPU, and now it&#8217;s more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.</p><p><strong>Akshat [00:24:01]:</strong> Yeah.</p><p><strong>Vibhu [00:24:02]:</strong> GPU, CPU. And now it&#8217;s like just constantly, and you just have to locate everything.</p><h2>Seventeen Clouds and the Supercloud Strategy</h2><p><strong>Akshat [00:24:08]:</strong> Yeah. And that&#8217;s one of the things that, again, we see as, something appealing about Modal, which is we&#8217;ve built this capacity pool that spans, 17 cloud providers, so we&#8217;re, we&#8217;re very good at Running on various kinds of cloud capacity across the world</p><p><strong>Swyx [00:24:24]:</strong> You don&#8217;t have your own data centers?</p><p><strong>Akshat [00:24:25]:</strong> We don&#8217;t have our own data centers. We just run across a lot of neo clouds</p><p><strong>Swyx [00:24:29]:</strong> Yeah. Are</p><p><strong>Akshat [00:24:30]:</strong> Metal providers.</p><p><strong>Swyx [00:24:30]:</strong> Yeah. Question mark.</p><p><strong>Swyx [00:24:31]:</strong> Yeah. You&#8217;re, you&#8217;re running the math, and you&#8217;re like, &#8220;What&#8217;s the cutover point where you&#8217;re like.&#8221;</p><p><strong>Akshat [00:24:36]:</strong> Yeah, it&#8217;s a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it&#8217;s worked out well because there are so many other people building data centers that we&#8217;re able to work effectively with them, and again, focus on what makes us, special.</p><p><strong>Swyx [00:24:55]:</strong> Yeah.</p><p><strong>Swyx [00:24:56]:</strong> 17 gets you into, like, the local providers sometimes. Like</p><p><strong>Akshat [00:25:00]:</strong> The,</p><p><strong>Swyx [00:25:01]:</strong> Which was the most interesting one?</p><p><strong>Akshat [00:25:02]:</strong> There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that&#8217;s why it&#8217;s something we&#8217;ve invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,</p><p><strong>Swyx [00:25:30]:</strong> Yeah</p><p><strong>Akshat [00:25:30]:</strong> You as a user would be able to.</p><p><strong>Swyx [00:25:32]:</strong> It&#8217;s a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.</p><p><strong>Akshat [00:25:41]:</strong> Yeah. That&#8217;s, that&#8217;s, that&#8217;s the idea. and so I guess when you mentioned colocation, that&#8217;s, that&#8217;s another interesting thing where, one thing we&#8217;ve seen is people come to us when they want, very specifically located, CPUs or GPUs, like they want</p><p><strong>Swyx [00:25:57]:</strong> Oh, they pin it in like</p><p><strong>Akshat [00:25:58]:</strong> Yeah</p><p><strong>Swyx [00:25:58]:</strong> EU?</p><p><strong>Akshat [00:25:59]:</strong> Exactly. Or EU, US.</p><p><strong>Swyx [00:26:01]:</strong> Right. Data resiliency</p><p><strong>Akshat [00:26:02]:</strong> Australia</p><p><strong>Swyx [00:26:02]:</strong> Locality thing or performance or what?</p><p><strong>Akshat [00:26:04]:</strong> It&#8217;s either data locality or latency, yeah.</p><p><strong>Swyx [00:26:07]:</strong> Yeah.</p><p><strong>Akshat [00:26:07]:</strong> Like, you want your. They&#8217;re running sandboxes and model. They want them to be right next to a</p><p><strong>Swyx [00:26:10]:</strong> Yeah, it&#8217;s easy then</p><p><strong>Akshat [00:26:11]:</strong> Yeah</p><p><strong>Swyx [00:26:12]:</strong> To. That is important in all those things. and so, like, you&#8217;ve accidentally, I don&#8217;t know if it&#8217;s accident, but, like, you&#8217;ve built the perfect primitive for agents to express themselves. And then, like, it&#8217;s almost very funny how every extra development just involves more file system, just involves more CPU.</p><p><strong>Akshat [00:26:30]:</strong> Yeah.</p><p><strong>Swyx [00:26:31]:</strong> Just like the things that you already have. I don&#8217;t know much about, if there&#8217;s any, like, networking usages that are interesting, but you&#8217;ve also done some good work on networking.</p><h2>Networking, Sidecars, Private IPv6, and Sandboxes</h2><p><strong>Akshat [00:26:40]:</strong> Yeah, that&#8217;s exactly right. Like, we&#8217;re just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.</p><p><strong>Swyx [00:26:49]:</strong> Yeah</p><p><strong>Akshat [00:26:50]:</strong> We see a few interesting networking things coming up. one is people want networked sandboxes. so we have</p><p><strong>Swyx [00:26:57]:</strong> For like a Docker cluster type thing.</p><p><strong>Akshat [00:26:59]:</strong> Yeah.</p><p><strong>Swyx [00:26:59]:</strong> Sorry, Docker Swarm. Oh, fuck. What is it called?</p><p><strong>Akshat [00:27:02]:</strong> Compose.</p><p><strong>Swyx [00:27:03]:</strong> Compose type thing.</p><p><strong>Akshat [00:27:04]:</strong> Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.</p><p><strong>Swyx [00:27:23]:</strong> Yeah.</p><p><strong>Akshat [00:27:23]:</strong> Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we&#8217;ve, we&#8217;ve had to build a lot of that stuff ourselves.</p><p><strong>Swyx [00:27:38]:</strong> Yeah.</p><p><strong>Akshat [00:27:39]:</strong> But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we&#8217;re seeing. We have support for that for a different reason, and yeah, we&#8217;ll see if that becomes stable.</p><p><strong>Swyx [00:27:52]:</strong> Like, just an open socket. It&#8217;s a. This is directly like mTLS.</p><p><strong>Akshat [00:27:56]:</strong> We do support that, which is you can, expose a tunnel inside a sandbox.</p><p><strong>Swyx [00:28:01]:</strong> Yeah.</p><p><strong>Akshat [00:28:01]:</strong> And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven&#8217;t talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.</p><p><strong>Akshat [00:28:28]:</strong> So it&#8217;s like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that&#8217;s truly serverless. and we did the overlay network for that. But then we&#8217;ve seen that people are using it for other reasons, and, I&#8217;m intrigued to yeah, what would people do with it.</p><p><strong>Swyx [00:28:59]:</strong> Build primitives and let people figure it out, right?</p><p><strong>Akshat [00:29:01]:</strong> Yeah, exactly.</p><p><strong>Swyx [00:29:02]:</strong> You put out a pretty interesting</p><p><strong>Akshat [00:29:03]:</strong> They&#8217;re like, they read the docs webpage. Let me use that</p><p><strong>Swyx [00:29:06]:</strong> Yeah</p><p><strong>Akshat [00:29:06]:</strong> Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they&#8217;re using it.</p><h2>RDMA, Memory Movement, and Distributed Training</h2><p><strong>Swyx [00:29:12]:</strong> Huh.</p><p><strong>Swyx [00:29:14]:</strong> The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I&#8217;m sure someone found it. It&#8217;s found it to be a lot more efficient before you made a thing out of it, right?</p><p><strong>Akshat [00:29:32]:</strong> Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.</p><p><strong>Akshat [00:29:39]:</strong> The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.</p><p><strong>Swyx [00:29:48]:</strong> Can I tell you, this is like a big aha moment for me because</p><p><strong>Akshat [00:29:51]:</strong> Yeah</p><p><strong>Swyx [00:29:51]:</strong> So I review 2,200 submissions for the World&#8217;s Fair.</p><p><strong>Akshat [00:29:56]:</strong> Yeah.</p><p><strong>Swyx [00:29:57]:</strong> And then I got this from John Osterhout</p><p><strong>Akshat [00:29:58]:</strong> Huh</p><p><strong>Swyx [00:29:59]:</strong> Who I don&#8217;t know if. Do John Osterhout by name?</p><p><strong>Akshat [00:30:01]:</strong> The name sounds familiar.</p><p><strong>Swyx [00:30:02]:</strong> He published a. He&#8217;s a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I&#8217;m like, you wouldn&#8217;t think that this guy, who is like operating systems guy, would care about RDMA.</p><p><strong>Akshat [00:30:20]:</strong> I, it makes sense to me because I,</p><p><strong>Swyx [00:30:24]:</strong> This is the cloud, right? Yeah</p><p><strong>Akshat [00:30:25]:</strong> Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there&#8217;s a lot of degrees of freedom, and it is a systems problem</p><p><strong>Swyx [00:30:41]:</strong> Yeah</p><p><strong>Akshat [00:30:41]:</strong> Moving memory around</p><p><strong>Swyx [00:30:42]:</strong> Yeah</p><p><strong>Akshat [00:30:43]:</strong> Scheduling.</p><p><strong>Swyx [00:30:44]:</strong> This shows you how primitive my understanding of networking stuff is.</p><p><strong>Swyx [00:30:46]:</strong> Is this like the domain of WireGuard as well?</p><p><strong>Akshat [00:30:50]:</strong> Not quite.</p><p><strong>Swyx [00:30:51]:</strong> It&#8217;s adjacent?</p><p><strong>Swyx [00:30:53]:</strong> Explain everything.</p><p><strong>Akshat [00:30:54]:</strong> Sure.</p><p><strong>Swyx [00:30:56]:</strong> How do we move memory around GPUs?</p><p><strong>Akshat [00:30:58]:</strong> Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you&#8217;ve set up.</p><p><strong>Swyx [00:31:09]:</strong> Yeah.</p><p><strong>Akshat [00:31:09]:</strong> Is it like it&#8217;s a VPN?</p><p><strong>Swyx [00:31:10]:</strong> Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you&#8217;re right. It is,</p><p><strong>Akshat [00:31:16]:</strong> Right. Yeah, you already moved on to new topics</p><p><strong>Swyx [00:31:17]:</strong> A similar</p><p><strong>Akshat [00:31:18]:</strong> Okay</p><p><strong>Swyx [00:31:19]:</strong> In the same space, WireGuard is, encrypted and this is,</p><p><strong>Akshat [00:31:23]:</strong> And you don&#8217;t need encryption.</p><p><strong>Swyx [00:31:23]:</strong> Yeah.</p><p><strong>Akshat [00:31:24]:</strong> Yeah.</p><p><strong>Swyx [00:31:24]:</strong> This is not encrypted. that&#8217;s the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you&#8217;re allowed to do it.</p><p><strong>Akshat [00:31:35]:</strong> Used to involve a full sidecar, but now you have eBPF in the Linux kernel.</p><p><strong>Swyx [00:31:39]:</strong> Yeah.</p><p><strong>Akshat [00:31:40]:</strong> Yeah. I don&#8217;t know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that&#8217;s the bottleneck, is your networking fast enough?</p><p><strong>Swyx [00:31:59]:</strong> Yeah. So I guess you&#8217;re talking about fully distributed training like, Dialog or something which is like cross data center</p><p><strong>Akshat [00:32:06]:</strong> That would be, yes.</p><p><strong>Swyx [00:32:07]:</strong> That&#8217;s the extreme.</p><p><strong>Akshat [00:32:08]:</strong> Yeah.</p><p><strong>Swyx [00:32:08]:</strong> You&#8217;re in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.</p><p><strong>Akshat [00:32:14]:</strong> When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it&#8217;s a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networking</p><p><strong>Swyx [00:32:40]:</strong> Okay</p><p><strong>Akshat [00:32:40]:</strong> Which is the standard that&#8217;s needed.</p><p><strong>Swyx [00:32:42]:</strong> Okay. So I misunderstood what</p><p><strong>Akshat [00:32:43]:</strong> 50</p><p><strong>Swyx [00:32:43]:</strong> What part of the stack you were</p><p><strong>Akshat [00:32:44]:</strong> 50 gigs over</p><p><strong>Swyx [00:32:45]:</strong> Yeah</p><p><strong>Akshat [00:32:45]:</strong> If you went</p><p><strong>Swyx [00:32:45]:</strong> Yeah</p><p><strong>Akshat [00:32:46]:</strong> RDMA.</p><p><strong>Swyx [00:32:46]:</strong> Okay.</p><p><strong>Swyx [00:32:48]:</strong> Yeah. I, very impressive work.</p><h2>Multi-Node Training, Post-Training, and Auto Research</h2><p><strong>Swyx [00:32:52]:</strong> So effectively you&#8217;re extending like the model philosophy to the training cluster, like, yeah.</p><p><strong>Akshat [00:32:59]:</strong> Yeah. And we&#8217;re, we&#8217;re not going for like large scale training runs. the thing that we&#8217;ve built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.</p><p><strong>Swyx [00:33:21]:</strong> Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.</p><p><strong>Akshat [00:33:31]:</strong> Yeah. The other use case we&#8217;ve seen for multi-node training is even if you have a big cluster, your researchers are still doing small runs</p><p><strong>Swyx [00:33:38]:</strong> Yes</p><p><strong>Akshat [00:33:39]:</strong> Having elasticity there</p><p><strong>Swyx [00:33:40]:</strong> Right, sure</p><p><strong>Akshat [00:33:40]:</strong> Matters a lot more.</p><p><strong>Swyx [00:33:41]:</strong> Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.</p><p><strong>Akshat [00:33:51]:</strong> We have a blog post on auto resource and model is,</p><p><strong>Swyx [00:33:55]:</strong> Yeah</p><p><strong>Akshat [00:33:56]:</strong> Yeah, like, turns out to be pretty good substrate for that.</p><p><strong>Swyx [00:33:59]:</strong> So my impression is auto research means many things, like</p><p><strong>Akshat [00:34:01]:</strong> Yeah</p><p><strong>Swyx [00:34:01]:</strong> Anything that Andrej coins. Right now it&#8217;s still science fair, right? Like not like, I don&#8217;t know how many people are doing this.</p><p><strong>Akshat [00:34:08]:</strong> We&#8217;re having a golf.</p><p><strong>Swyx [00:34:08]:</strong> Yeah.</p><p><strong>Akshat [00:34:09]:</strong> I thought the same thing.</p><p><strong>Swyx [00:34:11]:</strong> Yeah, you would know.</p><p><strong>Akshat [00:34:12]:</strong> We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we&#8217;ve automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It&#8217;ll even run like, NVIDIA inside profiler and it&#8217;ll like tweak configs and it&#8217;ll arrive the right thing. it&#8217;ll change your GPUs both from H200 to B200, and works really well.</p><p><strong>Swyx [00:34:47]:</strong> Nice.</p><p><strong>Akshat [00:34:47]:</strong> So yeah.</p><p><strong>Swyx [00:34:48]:</strong> By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.</p><p><strong>Swyx [00:34:52]:</strong> It&#8217;s very different from forward-deployed engineering from other people.</p><p><strong>Akshat [00:34:54]:</strong> Yeah. For our forward-deployed engineering team is, essentially they&#8217;re like applied inference researchers or applied training researchers.</p><p><strong>Swyx [00:35:02]:</strong> Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they&#8217;re good, they&#8217;re just like post-sale type of thing?</p><p><strong>Akshat [00:35:09]:</strong> It does, being able to talk to a customer and engage effectively with them</p><p><strong>Swyx [00:35:13]:</strong> Yeah</p><p><strong>Akshat [00:35:13]:</strong> Matters a lot.</p><p><strong>Swyx [00:35:14]:</strong> They want the same thing.</p><p><strong>Akshat [00:35:15]:</strong> Yeah.</p><p><strong>Swyx [00:35:15]:</strong> ?</p><p><strong>Akshat [00:35:15]:</strong> But it&#8217;s it&#8217;s not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.</p><p><strong>Swyx [00:35:23]:</strong> Okay. Let&#8217;s spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there&#8217;s all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.</p><h2>Auto Inference and Modal Bench</h2><p><strong>Akshat [00:35:51]:</strong> Yeah, like,</p><p><strong>Swyx [00:35:51]:</strong> Someone, some people call it like neural architecture search or whatever, right? Like.</p><p><strong>Akshat [00:35:54]:</strong> Yeah, - So the stuff I&#8217;ve seen people do with it is nowhere on the architecture level. It&#8217;s pretty much tweaking parameters, but it&#8217;s it&#8217;s a hyperparameter sweep that&#8217;s guided by some model intuition, so it&#8217;s like much more efficient than, whatever other, sweep you would have.</p><p><strong>Swyx [00:36:12]:</strong> Yeah, it&#8217;s just, it&#8217;s just a question of where you want to spend your compute?</p><p><strong>Akshat [00:36:16]:</strong> Right.</p><p><strong>Swyx [00:36:16]:</strong> &#8216;Cause yeah, you can just throw infinite amounts of money on this and somehow you&#8217;ll bang out Shakespeare?</p><p><strong>Akshat [00:36:22]:</strong> Yeah, infinite monkey.</p><p><strong>Swyx [00:36:24]:</strong> Yeah, so like the very good for model. and I think it&#8217;s also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.</p><p><strong>Akshat [00:36:42]:</strong> Yeah. They&#8217;re, they&#8217;re surprisingly good. I think like pre Cloud 4 they were not, and then now they&#8217;re able to shot, stuff out of the box. But we&#8217;re playing around with releasing like a Modal Bench for like the harder</p><p><strong>Swyx [00:36:55]:</strong> Yeah</p><p><strong>Akshat [00:36:55]:</strong> Things, that the LLMs cannot do yet and maybe</p><p><strong>Swyx [00:36:59]:</strong> What&#8217;s an example of that?</p><p><strong>Akshat [00:37:01]:</strong> I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It&#8217;s reasoning about that. But they&#8217;re able to shot, like</p><p><strong>Swyx [00:37:23]:</strong> Yeah. You can just add a skill to it?</p><h2>Compute Strategy and Capacity Planning</h2><p><strong>Akshat [00:37:26]:</strong> Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It&#8217;s to find things like that, so we can address them in our tool.</p><p><strong>Swyx [00:37:35]:</strong> Tune a skill. Yeah.</p><p><strong>Akshat [00:37:36]:</strong> Yeah.</p><p><strong>Swyx [00:37:36]:</strong> No. it&#8217;s it&#8217;s good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.</p><p><strong>Swyx [00:37:44]:</strong> Yeah.</p><p><strong>Akshat [00:37:45]:</strong> We have had a lot of growth, which means that, there&#8217;s - we&#8217;ve had to be much better about</p><p><strong>Swyx [00:37:53]:</strong> Planning</p><p><strong>Akshat [00:37:54]:</strong> Proactive capacity planning.</p><p><strong>Swyx [00:37:55]:</strong> Yeah.</p><p><strong>Akshat [00:37:55]:</strong> So we have,</p><p><strong>Swyx [00:37:57]:</strong> Which by the way, like it&#8217;s like a MBA&#8217;s like dream</p><p><strong>Akshat [00:38:00]:</strong> Yes</p><p><strong>Swyx [00:38:00]:</strong> Is like just planning this stuff. I think last time you and I talked about something maybe about this.</p><p><strong>Akshat [00:38:03]:</strong> Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on that</p><p><strong>Swyx [00:38:13]:</strong> Compute strategy?</p><p><strong>Akshat [00:38:13]:</strong> Yeah.</p><p><strong>Swyx [00:38:14]:</strong> I think,</p><p><strong>Akshat [00:38:14]:</strong> I feel like,</p><p><strong>Swyx [00:38:15]:</strong> I think the normies call it FP&amp;A or something.</p><p><strong>Akshat [00:38:18]:</strong> Well, it&#8217;s more It&#8217;s it&#8217;s not FP&amp;A. It&#8217;s it&#8217;s There&#8217;s a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,</p><p><strong>Swyx [00:38:49]:</strong> Yeah</p><p><strong>Akshat [00:38:49]:</strong> Based on that.</p><p><strong>Swyx [00:38:50]:</strong> Tokenomics.</p><p><strong>Akshat [00:38:50]:</strong> Yeah.</p><p><strong>Swyx [00:38:51]:</strong> This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.</p><p><strong>Akshat [00:38:59]:</strong> Yeah.</p><p><strong>Swyx [00:39:00]:</strong> And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost because</p><p><strong>Akshat [00:39:12]:</strong> Oh</p><p><strong>Swyx [00:39:12]:</strong> Compared to everyone else.</p><p><strong>Akshat [00:39:14]:</strong> Yeah. I hadn&#8217;t thought about that.</p><p><strong>Vibhu [00:39:16]:</strong> We&#8217;re at a fun time too?</p><p><strong>Akshat [00:39:18]:</strong> Yeah. It&#8217;s. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it&#8217;s how you can unlock more value for customers. Like, one of the things we&#8217;re building now is like a way for customers to get, If they don&#8217;t care about latency, like get much cheaper pricing and they&#8217;ll get results back in like next 24 hours or something, like a batch tier essentially.</p><h2>Batch Tiers and Latency-Insensitive Workloads</h2><p><strong>Swyx [00:39:47]:</strong> Yeah.</p><p><strong>Akshat [00:39:47]:</strong> And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficient</p><p><strong>Swyx [00:39:53]:</strong> Yeah. I feel like they&#8217;re not as popular. Like those, like the Frontier Labs have all those APIs. They&#8217;re not as popular as they should be.</p><p><strong>Akshat [00:40:00]:</strong> The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals and</p><p><strong>Swyx [00:40:08]:</strong> Okay</p><p><strong>Akshat [00:40:08]:</strong> Synthetic data prep and there it makes sense.</p><p><strong>Swyx [00:40:10]:</strong> Okay.</p><p><strong>Akshat [00:40:11]:</strong> But it&#8217;s from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don&#8217;t care about when they get it back.</p><p><strong>Swyx [00:40:22]:</strong> Yeah. And like they have a reasonable. It&#8217;s it&#8217;s also like a cousin to the stopping problem of like, will this finish in time?</p><p><strong>Akshat [00:40:30]:</strong> Yeah. You can bound it.</p><p><strong>Swyx [00:40:33]:</strong> Yeah.</p><p><strong>Akshat [00:40:33]:</strong> Like you can give people</p><p><strong>Swyx [00:40:34]:</strong> Yeah</p><p><strong>Akshat [00:40:34]:</strong> SLAs on it.</p><p><strong>Swyx [00:40:35]:</strong> Yeah. I think what&#8217;s, what&#8217;s interesting is like the next phase of model.</p><p><strong>Swyx [00:40:38]:</strong> Like what, do people expect from you, now that you&#8217;re established and you&#8217;re like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?</p><h2>What Modal Builds Next</h2><p><strong>Akshat [00:40:55]:</strong> We are building primitives that make our users&#8217; lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we&#8217;re thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven&#8217;t talked to anyone. It looks somewhat different on other verticals. Like, we&#8217;re also seeing a lot of real-time, audio-video stuff in there, which is why like, we&#8217;re working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it&#8217;s,</p><p><strong>Akshat [00:41:52]:</strong> We&#8217;re still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there&#8217;s a lot of other things people will need from this agent stack as they build production agents. So yeah, we&#8217;re thinking about those other things that fit in there.</p><p><strong>Swyx [00:42:13]:</strong> I want to ask what the other things are.</p><p><strong>Akshat [00:42:15]:</strong> Yeah. I probably should share right now.</p><p><strong>Swyx [00:42:17]:</strong> I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.</p><p><strong>Akshat [00:42:25]:</strong> Yeah.</p><p><strong>Swyx [00:42:25]:</strong> Because so far for me, it&#8217;s fine. so far for the. the first couple generations of cloud, it&#8217;s fine. What&#8217;s different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I&#8217;ll just kinda spew tokens at you until it like hopefully sparks something.</p><p><strong>Akshat [00:42:43]:</strong> Yeah.</p><p><strong>Swyx [00:42:44]:</strong> Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they&#8217;re like, &#8220;Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.&#8221; Is that it? like mediated permissions.</p><h2>Hard Guardrails vs. LLM-Mediated Permissions</h2><p><strong>Vibhu [00:43:03]:</strong> Now you&#8217;re looping it with a goal and letting it roll.</p><p><strong>Akshat [00:43:06]:</strong> Yeah, I&#8217;m, I&#8217;m skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.</p><p><strong>Swyx [00:43:16]:</strong> Yeah.</p><p><strong>Akshat [00:43:16]:</strong> Otherwise, someone can exfiltrate stuff.</p><p><strong>Swyx [00:43:20]:</strong> But like</p><p><strong>Akshat [00:43:20]:</strong> Yeah</p><p><strong>Swyx [00:43:20]:</strong> Maybe that&#8217;s old school thinking. Maybe we&#8217;re the dinosaurs.</p><p><strong>Swyx [00:43:23]:</strong> Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.</p><p><strong>Swyx [00:43:30]:</strong> Like it makes you feel uncomfortable.</p><p><strong>Akshat [00:43:31]:</strong> Yeah, I&#8217;m, I&#8217;m told</p><p><strong>Swyx [00:43:32]:</strong> But that&#8217;s what trusting the LLM is. Like imagine a spherical cow perfect LLM.</p><p><strong>Akshat [00:43:36]:</strong> Right.</p><p><strong>Swyx [00:43:37]:</strong> That it.</p><p><strong>Akshat [00:43:39]:</strong> Maybe.</p><p><strong>Swyx [00:43:41]:</strong> I wanna test the boundaries, right?</p><p><strong>Akshat [00:43:42]:</strong> Yeah.</p><p><strong>Swyx [00:43:42]:</strong> Like, and I don&#8217;t believe that, but I wanna see where I&#8217;m wrong &#8216;cause that&#8217;s, that&#8217;s the consensus.</p><p><strong>Akshat [00:43:49]:</strong> Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that&#8217;s gonna be a lot of mediated.</p><h2>Managed Agents and Specialized Sandboxes</h2><p><strong>Swyx [00:44:00]:</strong> There. I&#8217;ll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.</p><p><strong>Akshat [00:44:17]:</strong> Yeah.</p><p><strong>Swyx [00:44:17]:</strong> What&#8217;s going on?</p><p><strong>Akshat [00:44:19]:</strong> Yeah, we&#8217;re, very excited to partner with Anthropic and some of the other foundation labs, will not name who we&#8217;re also working with. the way we see it is the manage agent thing is a great place to start if you&#8217;re starting out building an agent and, But then when you get to, building something more production grade, like you&#8217;re a company that&#8217;s like Ramp that&#8217;s building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that&#8217;s the role that we are trying to play.</p><p><strong>Swyx [00:45:15]:</strong> Yeah</p><p><strong>Akshat [00:45:16]:</strong> We don&#8217;t really have an opinion on the harness, whether it runs - it&#8217;s a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We&#8217;ll see where people converge with that.</p><p><strong>Swyx [00:45:26]:</strong> Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?</p><p><strong>Akshat [00:45:31]:</strong> You mean like the OpenPipe</p><p><strong>Swyx [00:45:33]:</strong> OpenPipe is one. I think Vercel had one, which I can&#8217;t remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it&#8217;s kinda pseudo agent cloud type things.</p><p><strong>Akshat [00:45:50]:</strong> I personally have not played around with them.</p><p><strong>Swyx [00:45:53]:</strong> Yeah.</p><p><strong>Akshat [00:45:53]:</strong> Build agents with them.</p><p><strong>Swyx [00:45:54]:</strong> Everything&#8217;s bullish Modal, as long as it consumes more infra.</p><p><strong>Akshat [00:45:57]:</strong> That&#8217;s why we&#8217;re focusing on the infra layer. It&#8217;s somewhere where our, relative competence is and, also it&#8217;s a hard problem to solve.</p><p><strong>Swyx [00:46:06]:</strong> Yeah. I will say like just generally reflecting on that, I don&#8217;t know if - if there&#8217;s other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn&#8217;t really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, &#8220;Look at how many sandboxes I can spin up,&#8221; and no one gave a crap.</p><h2>Why Infrastructure Became Exciting Again</h2><p><strong>Akshat [00:46:39]:</strong> Yeah.</p><p><strong>Swyx [00:46:40]:</strong> And like now everyone gives a crap.</p><p><strong>Akshat [00:46:42]:</strong> That&#8217;s true. It is a very exciting time, and I think a lot of that&#8217;s driven by just the amount of scale all of this stuff needs.</p><p><strong>Swyx [00:46:50]:</strong> I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn&#8217;t necessarily have thought about it myself, which.</p><p><strong>Akshat [00:47:00]:</strong> We need the predictions.</p><p><strong>Swyx [00:47:02]:</strong> I think there&#8217;s a lot that you just don&#8217;t even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?</p><p><strong>Akshat [00:47:10]:</strong> What else is coming up for us</p><p><strong>Swyx [00:47:11]:</strong> Yeah. Where do you see things going?</p><p><strong>Akshat [00:47:13]:</strong> Yeah. I, in general</p><h2>Biotech, Robotics, and Non-LLM AI Workloads</h2><p><strong>Akshat [00:47:15]:</strong> It&#8217;s it&#8217;s clear that there&#8217;s there&#8217;s a huge shift happening. I think one thing that&#8217;s not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.</p><p><strong>Swyx [00:47:45]:</strong> Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?</p><p><strong>Akshat [00:47:50]:</strong> No. We,</p><p><strong>Swyx [00:47:51]:</strong> You should cloud only.</p><p><strong>Akshat [00:47:51]:</strong> Yeah.</p><p><strong>Swyx [00:47:52]:</strong> Yeah. Okay. But yeah, so what you&#8217;re saying is like because you&#8217;re focused on primitives and they&#8217;re good primitives, you find use cases in all these kinds of things.</p><p><strong>Akshat [00:48:01]:</strong> Yeah.</p><p><strong>Swyx [00:48:01]:</strong> Probably diversifies you a little bit away from LMS all the time.</p><p><strong>Akshat [00:48:05]:</strong> Yeah, absolutely. We&#8217;re, we&#8217;- our goal isn&#8217;t to only serve the LLM inference market.</p><p><strong>Swyx [00:48:10]:</strong> There are a lot just on the website, the audio,</p><p><strong>Akshat [00:48:12]:</strong> Yeah. We said both on</p><p><strong>Swyx [00:48:14]:</strong> Computational bio images. Yeah, there&#8217;s a lot here. There&#8217;s QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.</p><p><strong>Akshat [00:48:24]:</strong> Okay. Yeah.</p><p><strong>Swyx [00:48:25]:</strong> This screen reminds me of a fallen competitor, which Replicate.</p><h2>Model APIs vs. Differentiated AI Products</h2><p><strong>Swyx [00:48:31]:</strong> What&#8217;s your postmortem on what happened?</p><p><strong>Akshat [00:48:34]:</strong> This is one thing we&#8217;ve stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.</p><p><strong>Swyx [00:48:50]:</strong> Yeah.</p><p><strong>Akshat [00:48:50]:</strong> And we&#8217;ve always wanted to build for companies that are building products and need more flexibility that&#8217;s not just an API.</p><p><strong>Swyx [00:48:57]:</strong> Which you can build an API for a model and this is clearly what it is. But you - but what you&#8217;re saying, you can wrap it into a more fully functioning back end that you run.</p><p><strong>Akshat [00:49:06]:</strong> Yeah. So all of our examples, it&#8217;s not that spin up this model, here&#8217;s an API token, use it. They&#8217;re all code.</p><p><strong>Swyx [00:49:13]:</strong> Okay.</p><p><strong>Akshat [00:49:13]:</strong> And so the point is that this is just an example.</p><p><strong>Swyx [00:49:16]:</strong> Starter code.</p><p><strong>Akshat [00:49:17]:</strong> Yeah. But you can tweak it however you want.</p><p><strong>Swyx [00:49:20]:</strong> Yeah.</p><p><strong>Akshat [00:49:21]:</strong> And if you&#8217;re like a company building a product, like, computational bio whatnot, yeah.</p><p><strong>Swyx [00:49:26]:</strong> I guess I&#8217;m trying to tease out for listeners</p><p><strong>Akshat [00:49:28]:</strong> Yeah</p><p><strong>Swyx [00:49:28]:</strong> When does it stop becoming, oh, you&#8217;re just an API call and you&#8217;re just a wrapper on API to becoming what you call a product, right?</p><p><strong>Swyx [00:49:36]:</strong> Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?</p><p><strong>Akshat [00:49:46]:</strong> I think there&#8217;s a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that&#8217;s more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I&#8217;m not saying that&#8217;s, that successful, in that case. But a better example is like, let&#8217;s say Suno. because Suno, does not use Modal for training.</p><p><strong>Swyx [00:50:26]:</strong> Mikey on the pod. Yeah.</p><p><strong>Akshat [00:50:27]:</strong> But they use Modal for all their inference and that&#8217;s because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.</p><p><strong>Swyx [00:50:41]:</strong> It&#8217;s interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it&#8217;s a better model or agent that orchestrates video models.</p><h2>Video Agents and Production Workflows</h2><p><strong>Akshat [00:50:56]:</strong> Oh, interesting.</p><p><strong>Vibhu [00:50:56]:</strong> Language model backbone that can use tools</p><p><strong>Akshat [00:50:58]:</strong> Right</p><p><strong>Vibhu [00:50:59]:</strong> And write code.</p><p><strong>Akshat [00:51:00]:</strong> Like, yes, I can make my second video or my second video from Groq, but I want my minute video.</p><p><strong>Akshat [00:51:06]:</strong> And I&#8217;m not going there through normal video gen.</p><p><strong>Swyx [00:51:10]:</strong> Yeah, that&#8217;s interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,</p><p><strong>Akshat [00:51:22]:</strong> Yeah. Give it FFmpeg and just do it.</p><p><strong>Swyx [00:51:23]:</strong> Run FFmpeg. But like</p><p><strong>Akshat [00:51:25]:</strong> That&#8217;s not enough.</p><p><strong>Swyx [00:51:25]:</strong> Yeah.</p><p><strong>Akshat [00:51:26]:</strong> You need to give it Adobe.</p><p><strong>Swyx [00:51:27]:</strong> Yeah, I hadn&#8217;t put it together with like it would be a video production thing. in my mind these things were going more towards editing</p><p><strong>Akshat [00:51:36]:</strong> Yeah.</p><p><strong>Vibhu [00:51:36]:</strong> Well, shout out Mantis.</p><p><strong>Akshat [00:51:37]:</strong> I think about this a lot.</p><p><strong>Swyx [00:51:38]:</strong> .</p><p><strong>Akshat [00:51:41]:</strong> Yeah. Sorry.</p><p><strong>Vibhu [00:51:41]:</strong> Luma. Luma Agent is a version of this for video production, but it&#8217;s a off.</p><p><strong>Swyx [00:51:46]:</strong> I was gonna get your quick takes, on some other stuff that happens</p><h2>Gitpod/Ona, CI, and Runtime Sandboxes</h2><p><strong>Swyx [00:51:50]:</strong> In recent news and just-just see if you have anything interesting. Gitpod, very like-- somewhat like, different market. They&#8217;re in like the CI/CD market, but technically very impressive. I don&#8217;t know if you&#8217;ve like taken a real look at them.</p><p><strong>Akshat [00:52:03]:</strong> Yeah. we&#8217;ve, - People on our team have talked to the Gitpod team and they&#8217;- they&#8217;re technically very strong.</p><p><strong>Swyx [00:52:10]:</strong> Yeah.</p><p><strong>Akshat [00:52:10]:</strong> I - We&#8217;re, we&#8217;re very bullish at Modal on the CI market as well because</p><p><strong>Swyx [00:52:15]:</strong> Okay</p><p><strong>Akshat [00:52:15]:</strong> There&#8217;s, there&#8217;s more agents, coding agents.</p><p><strong>Swyx [00:52:18]:</strong> Yeah.</p><p><strong>Akshat [00:52:19]:</strong> They&#8217;re gonna run a lot more CI and the primitives there can be much better.</p><p><strong>Swyx [00:52:23]:</strong> I think there&#8217;s a lot of wasted CI.</p><p><strong>Akshat [00:52:25]:</strong> Yeah.</p><p><strong>Swyx [00:52:25]:</strong> So is it just like let&#8217;s filter? Like what is the highest order bid here in improving CI for agents?</p><p><strong>Akshat [00:52:32]:</strong> Well, there&#8217;s a lot of wasted time in CI on like</p><p><strong>Swyx [00:52:36]:</strong> Preparing</p><p><strong>Akshat [00:52:36]:</strong> Preparing your artifacts and like, getting you to the preparing your dependencies and whatnot.</p><p><strong>Swyx [00:52:44]:</strong> Oh.</p><p><strong>Akshat [00:52:44]:</strong> And, like build systems help with that. But like if you have primitives that are like memory snapshot and restore, can you just run CI more efficiently?</p><p><strong>Swyx [00:52:55]:</strong> Oh, okay. Okay. Okay. Interesting. Yeah. another form of like, demand compute.</p><p><strong>Akshat [00:53:02]:</strong> Yeah, exactly.</p><p><strong>Swyx [00:53:03]:</strong> Yeah.</p><p><strong>Akshat [00:53:03]:</strong> It needs the same again, platform.</p><p><strong>Swyx [00:53:06]:</strong> Yeah. So, for those who don&#8217;t know, Gitpod rebranded to Ona.</p><p><strong>Swyx [00:53:09]:</strong> It was like there was this whole thing. I - I like semi-sounded the alarm at Cognition. I was like, &#8220;You should take these guys seriously because their infra is very good.&#8221;</p><p><strong>Akshat [00:53:17]:</strong> Yeah.</p><p><strong>Swyx [00:53:18]:</strong> And but, then they join OpenAI and, presumably we&#8217;ll, we&#8217;ll see Codex Cloud from the Ona team.</p><p><strong>Swyx [00:53:26]:</strong> Like which I think would be very strong. - To me, like teams like that can set up the networking and like the secure boundaries for like, and your like agents to have their own cloud each, effectively is what you&#8217;re doing and I&#8217;m just trying to draw the analogy or the differences if you have studied them. Like what is the philosophical difference?</p><p><strong>Akshat [00:53:47]:</strong> My sense is maybe they didn&#8217;t go after the right market at the right time because - I guess also got lucky with like agent use cases really taking off and, needing, like more of like a sandbox shaped thing than like, my understanding is, yeah, Gitpod</p><p><strong>Swyx [00:54:06]:</strong> Really sandboxes work</p><p><strong>Akshat [00:54:07]:</strong> Never mind</p><p><strong>Swyx [00:54:07]:</strong> Like CI/</p><p><strong>Akshat [00:54:08]:</strong> Yeah</p><p><strong>Swyx [00:54:09]:</strong> Is sandboxes.</p><p><strong>Akshat [00:54:09]:</strong> Yeah.</p><p><strong>Swyx [00:54:10]:</strong> It&#8217;s just like build time sandboxes versus runtime sandboxes and it turned out runtime was better.</p><p><strong>Akshat [00:54:15]:</strong> Right. And the difference there is runtime sandboxes have a different configuration surface of like how you configure images, how you like attach like storage</p><p><strong>Swyx [00:54:25]:</strong> Yeah. It&#8217;s it&#8217;s fascinating. Other people, Astral also OpenAI.</p><h2>Python, TypeScript, and the Future of SDKs</h2><p><strong>Swyx [00:54:30]:</strong> Also like Python tooling ecosystem people. Are you still bullish build- building on top of Python? Also recently Modular also got bought by Qualcomm. Just any of your takes there?</p><p><strong>Akshat [00:54:43]:</strong> Yeah. we had Python as our first SDK language because that was the language that people did data and ML in. I now have Go and TypeScript SDKs as well. and our runtime is completely language- It is written in Rust, but it&#8217;s it&#8217;s not tied to Python by any means. We haven&#8217;t seen-- I think with like inference and training stuff, people are still very Python and the interesting thing with like the agent stuff is people use our TypeScript SDK a lot more because they&#8217;re not doing anything that needs ML.</p><p><strong>Akshat [00:55:13]:</strong> I don&#8217;t think we&#8217;ll have to go beyond that super soon</p><p><strong>Swyx [00:55:16]:</strong> Yeah</p><p><strong>Akshat [00:55:16]:</strong> &#8216;cause Python and TypeScript is still Dominant.</p><p><strong>Swyx [00:55:19]:</strong> The last two languages in the world.</p><p><strong>Akshat [00:55:21]:</strong> Yeah.</p><p><strong>Swyx [00:55:21]:</strong> That&#8217;s it.</p><p><strong>Akshat [00:55:22]:</strong> Well, English and prompting is the fourth language.</p><p><strong>Swyx [00:55:25]:</strong> English and prompting. I occasionally talk to people who try to build new languages. They&#8217;re like, - Even, what&#8217;s his face? Brett Taylor, who&#8217;s chairman of OpenAI was like, &#8220;We need a new language for LLMs.&#8221; So no one has come across one, and I keep looking. Python and TypeScript - You have a lot of data plus, but then also they are very imperfect as just as languages themselves. Then my close is, I think Modal used to be a big bet on developer experience.</p><h2>Agent Experience as a Company-Building Wedge</h2><p><strong>Swyx [00:55:52]:</strong> And you&#8217;ve pivoted the team to agent experience. Is it like the way now, like, do - do, - can entire companies and unicorns, multi-unicorns be built on just having better agent experience? Do you need something else?</p><p><strong>Akshat [00:56:05]:</strong> It&#8217;s a big part of our identity. it&#8217;s not just, like the very tactical, how does an agent use the CLI, but it&#8217;s also how easy is it to spin something up? Like, what is your iteration time when you wanna spin up a new service and, you wanna get something going in prod? in practice, that matters a lot, to people. And, I think it will continue to matter. Like, people are building stuff even faster, and if you give them ways to do it quickly not have overhead, then.</p><p><strong>Swyx [00:56:37]:</strong> I think the debate for me has been, do you do anything differently that is, like, very fundamentally different for developer experience versus agent experience?</p><p><strong>Swyx [00:56:44]:</strong> You seem to be on the side of they&#8217;re, they&#8217;re like this. They&#8217;re like cosine</p><p><strong>Akshat [00:56:48]:</strong> Yeah. We also have a blog post on that.</p><p><strong>Swyx [00:56:49]:</strong> Cosine similarity on, like, zero point nine or whatever.</p><p><strong>Akshat [00:56:53]:</strong> Yeah. pretty much it&#8217;s the main shift for us has been, as I said, like, we built this, benchmark, Modal Bench, to see where agents are lacking</p><p><strong>Swyx [00:57:02]:</strong> Yeah</p><p><strong>Akshat [00:57:02]:</strong> Literally add surface areas to a product if they&#8217;re reaching for something, like maybe this should just be a CLI.</p><p><strong>Swyx [00:57:09]:</strong> They halluc Oh, yeah. They hallucinate their own features.</p><p><strong>Akshat [00:57:11]:</strong> Yeah. And sometimes it makes sense. Like if they&#8217;re reaching for this thing, it&#8217;s product feedback. Like, give it to them. And then, yeah, moving-- we used to only have, like, logs and metrics in our UI, just moving all those things to the CLI as well, so they&#8217;re accessible in that form.</p><p><strong>Swyx [00:57:26]:</strong> Simple as that.</p><h2>Closing: Modal Bench, AX, and Execution</h2><p><strong>Swyx [00:57:28]:</strong> Cool. Thank you so much. Yeah.</p><p><strong>Akshat [00:57:29]:</strong> Yeah. Thank you.</p><p><strong>Swyx [00:57:30]:</strong> This was great.</p><p><strong>Akshat [00:57:30]:</strong> This was fun.</p><p><strong>Swyx [00:57:30]:</strong> Yeah. It was a great update and, I can see why you guys have succeeded so much. it is really, focus, but also really good execution.</p><p><strong>Akshat [00:57:39]:</strong> Thanks. we have a long way to go.</p><p><strong>Swyx [00:57:41]:</strong> All right. Thank you.</p><p><strong>Akshat [00:57:42]:</strong> Cool.</p>]]></content:encoded></item><item><title><![CDATA[[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI]]></title><description><![CDATA[a quiet day lets us read some condensed insight]]></description><link>https://www.latent.space/p/ainews-lilian-weng-summarizes-35</link><guid isPermaLink="false">https://www.latent.space/p/ainews-lilian-weng-summarizes-35</guid><pubDate>Wed, 08 Jul 2026 02:20:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!L_Ci!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Congrats to Meta Superintelligence on <a href="https://x.com/AIatMeta/status/2074577662840832382">having the top 2/3 image/video models</a> in the world! This would&#8217;ve been a candidate for a title story, but unfortunately that is pretty much all the detail we have about Muse Image/Video - no paper, no technical detail whatsoever. Still, this beats <a href="https://www.latent.space/p/ainews-microsoft-build-mai-thinking">the Microsoft MAI models from last month</a> which is nice.</p><p>We are noted <a href="https://news.smol.ai/issues?pattern=lilian%2520weng">Lilian Weng fans</a>, so we take notice whenever she drops another research recap, especially rare now that she is a cofounder at Thinky. Today she is thinking about the relationship of harnesses to RSI:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/lilianweng/status/2074372369213428144&quot;,&quot;full_text&quot;:&quot;new post on harness engineering for AI self-improvement: <a class=\&quot;tweet-url\&quot; href=\&quot;https://lilianweng.github.io/posts/2026-07-04-harness/\&quot;>lilianweng.github.io/posts/2026-07-&#8230;</a>\n\nIt is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter&quot;,&quot;username&quot;:&quot;lilianweng&quot;,&quot;name&quot;:&quot;Lilian Weng&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1923619459643711488/qmXOBhZ1_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-07T05:58:07.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:72,&quot;retweet_count&quot;:534,&quot;like_count&quot;:3888,&quot;impression_count&quot;:404538,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://lilianweng.github.io/posts/2026-07-04-harness/&quot;,&quot;title&quot;:&quot;Harness Engineering for Self-Improvement&quot;,&quot;description&quot;:&quot;The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an &#8220;ultraintelligent machine&#8221; as a system that can surpass humans in all intellectual activities and design better machines to improve itself. Yudkowsky (2008) used the phrase &#8220;recursive self-improvement&#8221; for a specific feedback loop: an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence. This feedback loop in modern AI may indicate the model rewriting its own weights directly, or more broadly the model improves the training pipeline and the deployment system, which in turn enables a better successor model with improved performance across economically valuable tasks. The speed of research development in AI has been shown to drastically accelerated in frontier labs (Anthropic; OpenAI).&quot;,&quot;domain&quot;:&quot;lilianweng.github.io&quot;,&quot;image&quot;:&quot;https://pbs.substack.com/news_img/2074372370534723585/sL23OMrz?format=png&amp;name=orig&quot;},&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>While we have written before about how <a href="https://www.latent.space/p/ainews-all-model-labs-are-now-agent?utm_source=publication-search">even Greg Brockman is now quietly endorsing agent/harness engineering</a>, it is refreshing for a respected thinker and neolab cofounder like Lilian to also agree that &#8220;<em>Even when many harness improvement[s] get eventually internalized into core model, <strong>the need to specify goals and context will not disappear</strong></em>.&#8221; </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BNEu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BNEu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 424w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 848w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 1272w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BNEu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png" width="1456" height="853" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:853,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:287124,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/205984146?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BNEu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 424w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 848w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 1272w, https://substackcdn.com/image/fetch/$s_!BNEu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5005c722-fdff-4ea8-aee0-6b37e44da978_1512x886.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><a href="https://lilianweng.github.io/posts/2026-07-04-harness/#harness-layer-vs-core-intelligence">Her post</a> breaks out the main proven design trends in harnesses that everyone should know, and then recaps the harness optimization literature, most notably from the well <a href="https://arxiv.org/abs/2510.04618">known ACE paper</a> to even more recent trends like <a href="https://arxiv.org/abs/2603.28052">Meta-Harnesses</a>,  which we have <a href="https://www.latent.space/p/ainews-its-meta-harness-summer">covered anecdotally on AINews</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!L_Ci!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!L_Ci!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 424w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 848w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 1272w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!L_Ci!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png" width="1456" height="1026" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1026,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:208013,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/205984146?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!L_Ci!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 424w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 848w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 1272w, https://substackcdn.com/image/fetch/$s_!L_Ci!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>It surely also provides a hint as to what Thinky is Thinking, beyond just <a href="https://www.latent.space/p/ainews-thinking-machines-native-interaction?utm_source=publication-search">Interaction Models</a>.</p><p></p><blockquote><p>AI News for 7/06/2026-7/07/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Agent Products, Harnesses, and Long-Running Workflows</strong></p><ul><li><p><strong>Anthropic expands &#8220;background agent&#8221; UX on top of Claude</strong>: The biggest product launch by engagement was <a href="https://x.com/claudeai/status/2074525815820169320">Claude Cowork coming to mobile and web</a>, positioning Claude as a task-running background teammate rather than a foreground chat UI. Related posts show the product convergence around a shared home tab and tighter Chat/Cowork integration from <a href="https://x.com/mikeyk/status/2074531605537046953">@mikeyk</a>. Separately, Anthropic extended access to <strong>Claude Fable 5</strong> on paid plans through July 12 in a highly engaged announcement from <a href="https://x.com/claudeai/status/2074548242386178258">@claudeai</a>, though many users noted the awkward timing relative to weekly limits in reactions from <a href="https://x.com/kimmonismus/status/2074606005963391225">@kimmonismus</a> and others.</p></li><li><p><strong>Harness engineering is increasingly the center of agent design</strong>: Lilian Weng&#8217;s new post was widely referenced as reframing recursive self-improvement around the <strong>harness</strong>, not direct weight self-modification; Sakana&#8217;s summary connects this to <strong>The AI Scientist</strong>, <strong>ShinkaEvolve</strong>, and <strong>Darwin G&#246;del Machine</strong> in <a href="https://x.com/SakanaAILabs/status/2074489949529776308">their thread</a>. LangChain echoed the same shift with a new <strong>Deep Agents</strong> course and an open-source harness project in posts from <a href="https://x.com/LangChain/status/2074539083204820997">@LangChain</a> and <a href="https://x.com/hwchase17/status/2074547871194698207">@hwchase17</a>. Google is also productizing this direction: Gemini API <strong>Managed Agents</strong> added <strong>background execution</strong>, <strong>remote MCP servers</strong>, <strong>custom function calling</strong>, and <strong>credential refresh</strong> in posts from <a href="https://x.com/_philschmid/status/2074533915038027972">@_philschmid</a> and <a href="https://x.com/OfficialLoganK/status/2074552932318765376">@OfficialLoganK</a>.</p></li><li><p><strong>Practical agent infra keeps getting more opinionated</strong>: There were several notable operator-facing updates: <strong>Codex Mobile iOS</strong> added task management, filtered diffs, SSH key login, branch comparison, and attachment flows in posts from <a href="https://x.com/Dimillian/status/2074396968223211819">@Dimillian</a> and <a href="https://x.com/reach_vb/status/2074400018769793176">@reach_vb</a>; <strong>Hermes Agent</strong> added pluggable secrets managers plus native <strong>1Password</strong> integration and export of sessions/datasets to formats including private Hugging Face repos in <a href="https://x.com/Teknium/status/2074564207555772912">@Teknium&#8217;s</a> <a href="https://x.com/Teknium/status/2074639961727655959">threads</a>; <strong>Weaviate 1.38</strong> made its MCP server GA with runtime-gated write access, notably allowing <strong>MCP_SERVER_WRITE_ACCESS_ENABLED</strong> to be flipped live without restart in <a href="https://x.com/victorialslocum/status/2074493681403339104">@victorialslocum&#8217;s post</a>. A more experimental pattern came from <a href="https://x.com/omarsar0/status/2074506169352180108">@omarsar0</a>, using a Dial MCP server so agents can escalate decisions via phone call/SMS/iMessage for human-in-the-loop control.</p></li></ul><p><strong>Model and Modality Releases: Audio, Speech, Robotics, and Media Generation</strong></p><ul><li><p><strong>Meta&#8217;s Muse Image/Muse Video push agentic generation into media</strong>: Meta Superintelligence Labs launched <strong>Muse Image</strong> and previewed <strong>Muse Video</strong> in announcements from <a href="https://x.com/AIatMeta/status/2074577662840832382">@AIatMeta</a>, <a href="https://x.com/alexandr_wang/status/2074555909347369105">@alexandr_wang</a>, and <a href="https://x.com/_tim_brooks/status/2074578008296628698">@_tim_brooks</a>. The notable technical angle is not just image quality, but an explicitly <strong>agentic generation loop</strong>: planning, web search, tool use, code execution, and self-refinement before rendering. Meta also says performance improves with <strong>scaled test-time compute</strong>, and that self-refinement behavior emerged during RL rather than being hand-scripted in <a href="https://x.com/AIatMeta/status/2074587864923250873">this follow-up</a>. On public evals, Muse Image quickly reached <strong>#2 on Image Arena</strong> behind GPT Image 2 in <a href="https://x.com/arena/status/2074581979765539153">Arena&#8217;s ranking</a>, while Muse Video debuted at <strong>#3 on Video Arena</strong> in <a href="https://x.com/arena/status/2074591193783320851">another Arena post</a>.</p></li><li><p><strong>NVIDIA and Cohere both shipped strong audio releases</strong>: NVIDIA released <strong>Audex</strong>, a <strong>30B parameter / 3B active MoE</strong> with <strong>1M context</strong> for unified text+audio work, summarized by <a href="https://x.com/HuggingPapers/status/2074384562952749254">@HuggingPapers</a> and described in more detail by <a href="https://x.com/_weiping/status/2074537900172050704">@_weiping</a>. The model&#8217;s core claim is preserving text intelligence while adding broad audio generation and understanding via a single MoE backbone. Cohere launched <strong>Cohere Transcribe Arabic</strong>, described as the most accurate open-source Arabic ASR model, under <strong>Apache 2.0</strong>, with emphasis on <strong>dialects</strong>, <strong>code-switching</strong>, and <strong>Arabic-accented English</strong> in posts from <a href="https://x.com/cohere/status/2074499759616729149">@cohere</a> and <a href="https://x.com/JayAlammar/status/2074511963934118282">@JayAlammar</a>.</p></li><li><p><strong>Open robotics keeps consolidating around Hugging Face + NVIDIA</strong>: NVIDIA expanded its robotics stack into the HF ecosystem by bringing <strong>GR00T 1.7</strong> and <strong>Isaac Teleop</strong> into <strong>LeRobot</strong>, aimed at open humanoid robotics workflows, in <a href="https://x.com/NVIDIARobotics/status/2074380795855147072">@NVIDIARobotics&#8217;s announcement</a> and <a href="https://x.com/NVIDIARobotics/status/2074390485251113317">integration guide</a>. On the embodied side, UMA showed a strong full-stack robotics narrative: <a href="https://x.com/RemiCadene/status/2074442725814878510">@RemiCadene</a> described a prototype built by a small team in 9 months, while <a href="https://x.com/RemiCadene/status/2074442439142609237">the Northstar reveal</a> and <a href="https://x.com/psermanet/status/2074512829617491996">@psermanet&#8217;s safety note</a> emphasized vertically integrated hardware/software for trustworthy robots.</p></li></ul><p><strong>Training, Inference, and Post-Training Techniques</strong></p><ul><li><p><strong>Liquid AI&#8217;s &#8220;Antidoom&#8221; directly targets reasoning-loop failure modes</strong>: One of the clearest technical releases of the day was <a href="https://x.com/liquidai/status/2074494130126811473">Liquid AI&#8217;s Antidoom</a>, an open-source training method to reduce <strong>doom loops</strong> where small reasoning models repeat tokens until context exhaustion. The reported reductions are substantial: <strong>LFM2.5-2.6B from 10.2% &#8594; 1.4%</strong> and <strong>Qwen3.5-4B from 22.9% &#8594; 1%</strong> under greedy sampling, with downstream eval gains. The method, <strong>FTPO (Final Token Preference Optimization)</strong>, relabels the loop-triggering token and redistributes probability toward alternatives, summarized well by <a href="https://x.com/helloiamleonie/status/2074498103982408044">@helloiamleonie</a> and <a href="https://x.com/LiorOnAI/status/2074547819114086561">@LiorOnAI</a>. This is a good example of the field&#8217;s recent pattern: removing specific failure modes rather than only scaling parameters.</p></li><li><p><strong>Inference efficiency and compression remain a major frontier</strong>: NVIDIA&#8217;s <strong>Puzzle-75B-A9B</strong> compression work got strong attention via <a href="https://x.com/omarsar0/status/2074543978129793462">@omarsar0</a>: compressing a hybrid MoE parent model while preserving reasoning, coding, long-context, and agentic quality, with roughly <strong>2x server throughput</strong> and <strong>1M-context concurrency on H100 rising from 1 request to 8</strong>. On the tooling side, <strong>Nsight Python 1.0</strong> launched in <a href="https://x.com/HagedornBastian/status/2074509770342445375">@HagedornBastian&#8217;s post</a>, making GPU perf analysis scriptable in Python. Unsloth also shipped <strong>GGUFs for DeepSeek-V4-Flash</strong>, plus export to <strong>NVFP4/FP8</strong> and speedups for <strong>GRPO</strong> and MoEs in <a href="https://x.com/danielhanchen/status/2074510444778463331">@danielhanchen&#8217;s update</a>.</p></li><li><p><strong>Agent RL and verification are getting more specialized</strong>: <a href="https://x.com/cwolferesearch/status/2074558199819067606">@cwolferesearch</a> highlighted how <strong>GRPO-style normalization</strong> is being adapted for agentic RL at the <strong>task</strong> or <strong>environment</strong> level to handle higher reward variance in multi-turn environments. Separately, <a href="https://x.com/omarsar0/status/2074556579580711050">@omarsar0</a> flagged a training-free <strong>verifier</strong> paper from Stanford/NVIDIA/Berkeley that reads calibrated continuous scores off scoring-token logits, posting strong numbers across <strong>Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench</strong> and suggesting verification is becoming an independent scaling axis.</p></li></ul><p><strong>Interpretability, Model Internals, and the &#8220;J-Space&#8221; Debate</strong></p><ul><li><p><strong>Anthropic&#8217;s J-space work dominated interpretability discussion, but also drew sharp criticism</strong>: The community split between seeing the work as useful mechanistic analysis and objecting to the consciousness framing. Strong critiques came from <a href="https://x.com/danburonline/status/2074429991576650014">@danburonline</a>, <a href="https://x.com/paul_cal/status/2074388528243310976">@paul_cal</a>, and <a href="https://x.com/scaling01/status/2074432865794679235">@scaling01</a>, who argued the vectors are causal largely by construction under the Jacobian-lens definition. A useful historical reference came from <a href="https://x.com/jacobandreas/status/2074487546692735002">@jacobandreas</a>, pointing readers back to the original <strong>Jacobian lenses</strong> paper.</p></li><li><p><strong>The stronger technical takeaway is cross-model structure, not consciousness rhetoric</strong>: <a href="https://x.com/eliebakouch/status/2074532904009421260">@eliebakouch</a> computed <strong>CKA similarity</strong> on J-lens geometry across <strong>38 open models</strong> and found surprisingly universal layer/depth organization, even across unrelated families like <strong>Llama</strong> and <strong>OLMo</strong>. Anthropic and Neuronpedia also released <strong>J-lens weights for open models</strong>, noted in <a href="https://x.com/eliebakouch/status/2074537985102565795">this follow-up</a>. In parallel, Goodfire introduced <strong>Block-Sparse Featurizers</strong> for multidimensional concepts in activations, arguing many vision concepts are inherently <strong>2&#8211;4 dimensional blocks</strong> rather than single directions, in <a href="https://x.com/GoodfireAI/status/2074634702737281303">their thread</a>.</p></li></ul><p><strong>Benchmarks, Evaluations, and Domain-Specific Systems</strong></p><ul><li><p><strong>Agent and legal benchmarks continue to expose the gap between &#8220;passes many criteria&#8221; and &#8220;fully solves real work&#8221;</strong>: <a href="https://x.com/arena/status/2074484787663052849">Agent Arena</a> placed <strong>Claude Sonnet 5 (Thinking)</strong> at <strong>#6</strong>, with strongest signals in confirmed task success and bash usage, but still with uncertainty around steerability. Artificial Analysis launched <strong>Harvey LAB-AA</strong>, a legal-agent benchmark over <strong>120 private legal tasks across 24 practice areas</strong>, where <strong>Claude Fable 5</strong> led at <strong>14.2% all-pass rate</strong>; <strong>Claude Opus 4.8</strong> and <strong>GLM-5.2</strong> tied at <strong>7.5%</strong>, with GLM hitting that at roughly <strong>~6% of Fable&#8217;s cost per task</strong> in <a href="https://x.com/ArtificialAnlys/status/2074541975186165887">their release</a>. The big message is that models can satisfy many individual rubric items yet still fail to produce acceptable end-to-end deliverables.</p></li><li><p><strong>Research automation and specialized domain systems are broadening</strong>: Google promoted <strong>Experience AI Scientist</strong>, a multi-agent system for end-to-end scientific workflows, in <a href="https://x.com/GoogleResearch/status/2074384746076135575">this ICML post</a>. DeepMind also launched <strong>Predicting the Past</strong>, grounding Gemini in <strong>Aeneas</strong> and <strong>Ithaca</strong> for Greek/Latin historical analysis via plain-English interactions, in <a href="https://x.com/GoogleDeepMind/status/2074513661750546762">their thread</a>. On legal AI commercialization, <strong>Norm Ai</strong> announced a <strong>$120M Series C at $1.2B valuation</strong> and described a full-stack &#8220;agentic law&#8221; setup spanning software plus an AI-native law firm in <a href="https://x.com/johnjnay/status/2074485345593245833">@johnjnay&#8217;s post</a>.</p></li></ul><p><strong>Top tweets (by engagement)</strong></p><ul><li><p><strong>Claude access / product rollout</strong>: <a href="https://x.com/claudeai/status/2074525815820169320">Claude Cowork on mobile and web</a> and <a href="https://x.com/claudeai/status/2074548242386178258">Fable 5 access extended through July 12</a> were the most-engaged technically relevant product announcements.</p></li><li><p><strong>Open-source developer program</strong>: <a href="https://x.com/ClaudeDevs/status/2074570404035993780">@ClaudeDevs offering 6 months of Claude Max 20x for open-source maintainers</a> drew massive engagement and is likely to matter for tool adoption in OSS ecosystems.</p></li><li><p><strong>Meta media generation</strong>: <a href="https://x.com/AIatMeta/status/2074577662840832382">Muse Image launch</a> and <a href="https://x.com/arena/status/2074581979765539153">Arena&#8217;s #2 ranking for Muse Image</a> were the biggest multimodal product stories.</p></li><li><p><strong>Reasoning reliability</strong>: <a href="https://x.com/liquidai/status/2074494130126811473">Liquid AI&#8217;s Antidoom release</a> stood out as the day&#8217;s highest-signal training technique post.</p></li><li><p><strong>Interpretability</strong>: <a href="https://x.com/eliebakouch/status/2074532904009421260">Cross-model J-lens universality across 38 open models</a> was the strongest technical follow-on to the J-space discourse.</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. Open Model Releases and Inference Efficiency</strong></h3><ul><li><p><strong><a href="https://www.reddit.com/r/LocalLLaMA/comments/1uoozt4/new_open_model_from_tencent_hy_hy3_295b_total_21b/">New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)</a></strong> (Activity: 653): <strong>Tencent released the non-preview Hy3 open model collection on <a href="https://huggingface.co/collections/tencent/hy3">Hugging Face</a>, described as a </strong><code>295B</code><strong>-parameter MoE with </strong><code>21B</code><strong> active parameters, now under Apache 2.0 rather than the prior restrictive community license. The post highlights that the earlier license reportedly excluded use in regions including South Korea, the UK, and the EU, while top comments point to claimed benchmark gains over HY3-Preview and frame this as potentially relevant for high-end local/home inference setups.</strong> Commenters viewed the Apache 2.0 relicensing as the most important change, especially given Tencent&#8217;s recent translation models also using Apache licensing. There was cautious optimism that the reported benchmark improvements may translate to real-world usefulness, but with implicit skepticism until tested outside vendor charts.</p><ul><li><p>Commenters highlighted that <strong>Hunyuan/HY3</strong> is now listed as <strong>Apache 2.0</strong>, contrasting it with the prior &#8220;community&#8221; license that reportedly restricted usage in regions such as <strong>South Korea, the UK, and the EU</strong>. This was viewed as technically important for deployment because Apache 2.0 removes many commercial and geographic usage barriers.</p></li><li><p>Several users focused on whether Tencent&#8217;s claimed benchmark improvements over <strong>HY3-Preview</strong> will translate into real-world workloads. Given the reported <code>295B</code><strong> total / </strong><code>21B</code><strong> active</strong> MoE-style configuration, commenters suggested it could be relevant for &#8220;high-end home setups&#8221; if inference formats such as <strong>GGUF</strong> become available.</p></li><li><p>There was early speculation that HY3 could become an alternative to <strong>Qwen</strong> and <strong>MiniMax</strong> models in local/open-weight workflows, but commenters were waiting for quantized releases and independent testing before drawing conclusions.</p></li></ul></li></ul><p></p><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-lilian-weng-summarizes-35">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[[AINews] The Field Guide to Fable]]></title><description><![CDATA[a quiet day lets us digest the world's most significant model launch... to date.]]></description><link>https://www.latent.space/p/ainews-the-field-guide-to-fable</link><guid isPermaLink="false">https://www.latent.space/p/ainews-the-field-guide-to-fable</guid><pubDate>Tue, 07 Jul 2026 04:44:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/9fubhllmsBU" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>While we congratulate <a href="https://www.latent.space/p/world-models-and-general-intuition">(friend of the show!) General Intuition</a> on <a href="https://x.com/gen_intuition/status/2074104524596457706">their new model</a> and <a href="https://www.latent.space/p/shunyu">(friend of the show!) Shunyu Yao </a>on <a href="https://x.com/ShunyuYao12/status/2074151389945827744">their new model</a>, and the world awaits the release of <a href="https://news.ycombinator.com/item?id=48799614">GPT-5.6 Sol Ultra</a>, people are racing to find the limits of Fable 5 before the <a href="https://www.latent.space/p/ainews-sonnet-5-today-and-fable-5">subscription subsidy ends tomorrow</a>. </p><p>Thariq had been working on a &#8220;<a href="https://x.com/trq212/status/2073100352921215386">Field Guide to Fable</a>&#8221; blog series, and happened to have a keynote planned the day of the relaunch, so he kindly pivoted the entire keynote in one night to give the most timely advice he had, which was released today:</p><div id="youtube2-9fubhllmsBU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;9fubhllmsBU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/9fubhllmsBU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>The 4 segments are (my watchalong commentary in italics):</p><ul><li><p><a href="https://www.youtube.com/watch?v=9fubhllmsBU"><span>0:00</span></a><span> Introduction and setting the stage for Fable</span></p></li><li><p><a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=152s"><span>2:32</span></a><span> </span><strong><span>Unhobbling Claude:</span></strong><span> Understanding model behavior</span></p><ul><li><p><em>The constraints on a model are often imposed by US - </em><strong>&#8220;the harness we put them in, and the way we prompt them&#8221;</strong><em>. Therefore when we encounter a new class of model, we should expect to remove or change those harnesses and prompts in order to elicit new behaviors that you otherwise would never see because you were overly limiting (aka hobbling) the model.</em></p></li><li><p><em>Case in point: most people have come to agree with Thariq on the <a href="https://x.com/trq212/status/2052809885763747935">unreasonable effectiveness of HTML</a>.</em></p></li></ul></li><li><p><a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=548s"><span>9:08</span></a><span> </span><strong><span>Finding your unknowns</span></strong><span>: Navigating the gap between map and territory</span></p><ul><li><p><a href="https://x.com/trq212/status/2073100352921215386">already blogged here</a>.</p></li><li><p><em>a close cousin to &#8220;unhobbling&#8221; - if unhobbling is about clearing out outdated knowns, then this is about finding things you didn&#8217;t even know you didn&#8217;t know.</em></p></li><li><p><em>easiest techniques:</em></p><ul><li><p><em>telling claude to do a &#8220;<strong>blindspot pass</strong>&#8221; for your unknowns</em></p></li><li><p><em><strong>brainstorm</strong> for &#8220;wildly different design directions&#8221;</em></p></li><li><p><em><strong>interview me</strong> - similar to <a href="https://www.youtube.com/watch?v=v4F1gFy-hqg&amp;t=132s">/grill-me</a>, but prioritizing high impact questions</em></p></li><li><p><em><strong>use references</strong>: in the case of migrations</em></p></li><li><p><em><strong>keep implementation-notes.md</strong>: a running log of underspecified decisions made on your behalf</em></p></li><li><p><em><strong>quiz me</strong> - ensure MY understanding</em></p></li></ul></li></ul></li><li><p><a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=869s"><span>14:29</span></a><span> </span><strong><span>Dealing with Grief: </span></strong><span>Reflecting on the emotional shift in coding productivity</span></p><ul><li><p><em>What you used to spend weeks on is now done in hours</em></p></li></ul></li><li><p><a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=990s"><span>16:30</span></a><span> </span><strong><span>Being unreasonable</span></strong><span>: Demanding good, fast, and cheap results</span></p><ul><li><p>&#8220;<strong>Tradeoffs are not real</strong>&#8221;<em> - because Fable is more capable, you can be more ambitious and not accept tradeoffs.</em></p></li><li><p>&#8220;<em>Building is easy, generating value is still hard&#8221;</em>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!p3LG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!p3LG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 424w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 848w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 1272w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!p3LG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png" width="862" height="1396" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1396,&quot;width&quot;:862,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1621224,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/205713711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!p3LG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 424w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 848w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 1272w, https://substackcdn.com/image/fetch/$s_!p3LG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179f6ab6-62ad-492e-bb2c-67f7f2dbb861_862x1396.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p></li></ul></li></ul><p>Overall, an excellent talk that we will be mapping out the implications of as the world acclimatizes to the first Fable-class models.</p><p></p><blockquote><p>AI News for 7/04/2026-7/06/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Tencent Hunyuan&#8217;s Hy3 Release and the Open-Weight Frontier</strong></p><ul><li><p><strong>Hy3 lands as a serious open model</strong>: Tencent released <strong>Hy3</strong> under <strong>Apache 2.0</strong>, a <strong>295B MoE</strong> with <strong>21B active parameters</strong>, <strong>192 experts / top-8 routing</strong>, <strong>GQA</strong>, <strong>256K context</strong>, and a <strong>3.8B MTP layer</strong> for speculative decoding. Multiple posts framed it as competitive with much larger systems on reasoning, coding, and agentic tasks, with particular emphasis on reliability improvements like tool-calling stability and anti-hallucination work <a href="https://x.com/eliebakouch/status/2074011171661701466">@eliebakouch</a>, <a href="https://x.com/HuggingPapers/status/2074024501201813797">@HuggingPapers</a>, <a href="https://x.com/ShunyuYao12/status/2074151389945827744">@ShunyuYao12</a>.</p></li><li><p><strong>Inference support was unusually day-0 mature</strong>: <a href="https://x.com/vllm_project/status/2074147504254517529">@vllm_project</a> said Hy3 runs natively in <strong>vLLM</strong> from launch with tool-call and reasoning parsers, <strong>MTP speculative decoding</strong>, and validated support on <strong>NVIDIA and AMD</strong>. A follow-up detailed Tencent production kernels now upstreamed into vLLM main, including load-balanced decode scheduling and fused FP8 MoE serving, with reported gains of <strong>up to 2.95x</strong> on mixed-length decode and latency reductions of roughly <strong>24% TTFT</strong> and <strong>17% TPOT</strong> versus default backends <a href="https://x.com/vllm_project/status/2074147506875969754">@vllm_project</a>. Community reaction was strong enough that <a href="https://x.com/Teknium/status/2074264567803531589">@Teknium</a> quickly made Hy3 free on Nous Portal for two weeks.</p></li><li><p><strong>Broader open-model context</strong>: Hy3 was immediately compared against <strong>GLM-5.2</strong>, with some posters arguing Tencent has now joined the very top tier of open-source labs if the benchmark and vibe-test results hold <a href="https://x.com/teortaxesTex/status/2074012467886178725">@teortaxesTex</a>, while others still maintained <strong>GLM-5.2</strong> as the best currently usable open-weight model in practice <a href="https://x.com/__tinygrad__/status/2074206866641752190">@</a><strong><a href="https://x.com/__tinygrad__/status/2074206866641752190">tinygrad</a></strong>, <a href="https://x.com/mbusigin/status/2074238100251799998">@mbusigin</a>. The net takeaway: the open frontier is compressing fast, and the competition is increasingly about deployment robustness rather than just raw leaderboard deltas.</p></li></ul><p><strong>Agent Benchmarks, Harnesses, and Long-Running Memory</strong></p><ul><li><p><strong>AutomationBench-AA adds a more realistic agent eval</strong>: <a href="https://x.com/ArtificialAnlys/status/2074194764510208230">@ArtificialAnlys</a> launched an independent leaderboard for Zapier&#8217;s <strong>AutomationBench</strong>, evaluating agents across <strong>657 tasks</strong> and <strong>40 simulated SaaS apps</strong> with both objectives and guardrails. <strong>Claude Fable 5</strong> led at <strong>48.6%</strong>, narrowly ahead of <strong>Opus 4.8</strong> at <strong>48.5%</strong>, with <strong>Gemini 3.5 Flash</strong> at <strong>42.6%</strong> and <strong>GPT-5.5 xhigh</strong> at <strong>42.1%</strong>. More interesting than the ranking: every model still breaks business rules, and Gemini looked notably strong on <strong>objective-per-guardrail-violation</strong> and <strong>cost efficiency</strong>. Open weights remain meaningfully behind, with <strong>GLM-5.2 max</strong> the best listed open model at <strong>27.8%</strong>.</p></li><li><p><strong>Capability indices are becoming multidimensional</strong>: Artificial Analysis also introduced six domain-specific indices&#8212;<strong>Finance &amp; Accounting, Legal, Healthcare &amp; Medical, Strategy &amp; Ops, Engineering, Economics</strong>&#8212;to move past single scalar model scores <a href="https://x.com/ArtificialAnlys/status/2074299714699469221">@ArtificialAnlys</a>. The headline was familiar&#8212;<strong>Claude Fable 5</strong> plus <strong>Opus 4.8 fallback</strong> leads&#8212;but the more useful insight is how sharply rankings reshuffle by domain and how steep the price/performance frontier has become. This aligns with <a href="https://x.com/fchollet/status/2074242671103889799">@fchollet</a>, who argued that reporting benchmark scores without <strong>cost per task</strong> is increasingly meaningless.</p></li><li><p><strong>Memory and retrieval remain bottlenecks for persistent agents</strong>: Two papers got traction here. First, <strong>A-TMA</strong> tackles &#8220;ghost memory,&#8221; where stale and current facts are retrieved together in long-running assistants; on the LTP benchmark, adding it to Graphiti reportedly improves conflict accuracy by <strong>+0.240 absolute</strong> <a href="https://x.com/omarsar0/status/2074121191846261022">@omarsar0</a>. Second, <strong>ReContext</strong> is a training-free long-context inference harness that replays model-internal evidence right before answer generation, improving evidence utilization across eight 128K datasets <a href="https://x.com/dair_ai/status/2074178316819677238">@dair_ai</a>. Combined with <strong>BlockSearch</strong> for million-token in-context retrieval <a href="https://x.com/dair_ai/status/2074117920133898707">@dair_ai</a>, the theme is clear: better memory behavior is increasingly being engineered at inference time, not just trained in.</p></li></ul><p><strong>Anthropic&#8217;s J-Space / Global Workspace Results</strong></p><ul><li><p><strong>Mechanistic interpretability took center stage</strong>: Anthropic released research claiming a <strong>global-workspace-like internal structure</strong> in Claude, centered on a small subset of activations they call <strong>J-space</strong> <a href="https://x.com/AnthropicAI/status/2074185348142280912">@AnthropicAI</a>, <a href="https://x.com/AnthropicAI/status/2074185387577094398">@AnthropicAI</a>. The core claim is not chain-of-thought extraction, but identification of a privileged internal representational substrate that appears available for report, modulation, and flexible reasoning. Anthropic also shipped a Neuronpedia demo for open-weight models <a href="https://x.com/AnthropicAI/status/2074185390060110138">@AnthropicAI</a>.</p></li><li><p><strong>Why researchers cared</strong>: Interpretability researchers treated this as stronger evidence for a model &#8220;working memory&#8221; or internal workspace than prior public work, even if they disagreed with the framing. <a href="https://x.com/NeelNanda5/status/2074193936588148891">@NeelNanda5</a> called it the best evidence yet for a working-memory-like mechanism. <a href="https://x.com/Jack_W_Lindsey/status/2074215950602379388">@Jack_W_Lindsey</a> argued understanding this privileged space could be key to LLM cognition. Posts also highlighted practical safety angles: the workspace can reportedly surface hidden concepts, detect prompt injections, and expose internal sabotage-related features before they are verbalized <a href="https://x.com/mlpowered/status/2074190714100146483">@mlpowered</a>, <a href="https://x.com/LiorOnAI/status/2074198891990548940">@LiorOnAI</a>, <a href="https://x.com/omarsar0/status/2074264122330612223">@omarsar0</a>.</p></li><li><p><strong>But the &#8220;consciousness&#8221; language was contested</strong>: Anthropic&#8217;s public framing invited strong pushback. Supporters said the results suggest a functional analog of <strong>access consciousness</strong> rather than phenomenal consciousness <a href="https://x.com/BorisMPower/status/2074201312531734567">@BorisMPower</a>, while critics argued the company was overclaiming by conflating privileged latent activation with consciousness <a href="https://x.com/AlanCowen/status/2074265992570736919">@AlanCowen</a>. Even some sympathetic takes emphasized the bigger story is a new <strong>intervention point</strong> for auditing and steering models, not philosophy.</p></li></ul><p><strong>Inference, Serving, and Systems Efficiency</strong></p><ul><li><p><strong>Speculative decoding remains hot infrastructure</strong>: <a href="https://x.com/lmsysorg/status/2074176669108367549">@lmsysorg</a> added <strong>DSpark</strong> to SGLang for confidence-driven, variable-length verification. The pitch is that under high load it avoids verifying every draft token, improving the throughput/latency tradeoff relative to fixed-budget speculative methods; DeepSeek-V4-Pro reportedly reached <strong>383.7 tok/s at batch=1 on B300</strong>. Microsoft also discussed prompt-level optimization of <strong>GPT-5.5</strong> in the GitHub Copilot harness to improve latency and token efficiency after launch <a href="https://x.com/code/status/2074178799512539571">@code</a>, <a href="https://x.com/pierceboggan/status/2074180737147027757">@pierceboggan</a>.</p></li><li><p><strong>Inference efficiency is increasingly the strategic bottleneck</strong>: <a href="https://x.com/jon_durbin/status/2074169183835685351">@jon_durbin</a> argued that inference, not training alone, is now &#8220;the whole game,&#8221; because every data pipeline, RL loop, and agent runtime ultimately cashes out as test-time compute. That perspective also showed up in lower-level kernel work: Chutes reported major speedups for <strong>MiniMax MSA</strong> and <strong>GatedDeltaNet-2</strong>, including <strong>~7x</strong> sparse-attention training improvements on <strong>RTX Pro 6000 / SM120</strong> and better fused FP8 kernels <a href="https://x.com/jon_durbin/status/2074119835366134188">@jon_durbin</a>.</p></li><li><p><strong>Infra releases beyond model serving</strong>: Cloudflare launched <strong>Workers Cache</strong>, a regionally tiered cache in front of Worker entrypoints configured via standard HTTP headers <a href="https://x.com/Cloudflare/status/2074117419728007181">@Cloudflare</a>. OpenAI shipped <strong>GPT-Realtime-2.1-mini</strong>, bringing reasoning and tool use to the mini realtime line at the same price as the prior mini, alongside claimed <strong>25%+ p95 latency reductions</strong> from caching improvements <a href="https://x.com/OpenAIDevs/status/2074255408013955466">@OpenAIDevs</a>, <a href="https://x.com/OpenAIDevs/status/2074255420831735824">@OpenAIDevs</a>.</p></li></ul><p><strong>World Models, Speech, and Document AI</strong></p><ul><li><p><strong>MIRA is a notable world-model demo</strong>: General Intuition and Kyutai, with Epic Games, introduced <strong>MIRA</strong>, a playable multiplayer world model for Rocket League trained on <strong>10k hours</strong> of bot-collected data <a href="https://x.com/gen_intuition/status/2074104524596457706">@gen_intuition</a>. It runs in real time at <strong>20 fps</strong>, and posts highlighted a <strong>5B-parameter</strong> model running an entire 2v2 match on a single <strong>NVIDIA B200</strong>, with no explicit physics or rendering engine <a href="https://x.com/TheRundownAI/status/2074184559768277398">@TheRundownAI</a>. This was one of the clearest signals that video/world-model work is moving from toy demos toward interactive simulators.</p></li><li><p><strong>Speech remains highly competitive</strong>: AssemblyAI released <strong>Universal-3.5 Pro Realtime</strong>, a streaming STT model with <strong>4.1% WER</strong> on AA-WER Streaming and contextual priming that can be updated mid-call without reconnecting <a href="https://x.com/ArtificialAnlys/status/2074160133702402314">@ArtificialAnlys</a>. On the TTS side, Artificial Analysis said <strong>Speechify Simba 3.2</strong> now leads its Speech Arena at <strong>1233 Elo</strong>, ahead of Gemini 3.1 Flash TTS, Sonic 3.5, and Inworld Realtime TTS 1.5 Max, while also being the cheapest among top-ranked models <a href="https://x.com/ArtificialAnlys/status/2074265309985570890">@ArtificialAnlys</a>.</p></li><li><p><strong>Document-context pipelines are becoming multimodal by default</strong>: LlamaIndex and LanceDB described a retrieval pipeline for messy PDFs that separates <strong>pages, chunks, and extracted assets</strong> into linked multimodal tables, reporting <strong>82% any-page-hit@5</strong> and <strong>74% answer accuracy</strong> on a labeled ESG-report benchmark <a href="https://x.com/lancedb/status/2074153945631457663">@lancedb</a>, <a href="https://x.com/llama_index/status/2074170470119752084">@llama_index</a>. This pairs with Jerry Liu&#8217;s broader argument for a dedicated &#8220;document context layer&#8221; for agents <a href="https://x.com/jerryjliu0/status/2074165277634253106">@jerryjliu0</a>.</p></li></ul><p><strong>Top tweets (by engagement)</strong></p><ul><li><p><strong>Anthropic&#8217;s global workspace paper</strong> dominated engagement, with the primary announcement on Claude&#8217;s internal workspace/J-space far above everything else <a href="https://x.com/AnthropicAI/status/2074185348142280912">@AnthropicAI</a>.</p></li><li><p><strong>Tencent Hy3</strong> was the biggest pure model-release story, especially among technical accounts discussing open-source competitiveness and deployment <a href="https://x.com/teortaxesTex/status/2074012467886178725">@teortaxesTex</a>, <a href="https://x.com/ShunyuYao12/status/2074151389945827744">@ShunyuYao12</a>.</p></li><li><p><strong>MIRA&#8217;s playable world model</strong> was the standout multimodal/system demo <a href="https://x.com/gen_intuition/status/2074104524596457706">@gen_intuition</a>.</p></li><li><p><strong>Will Depue&#8217;s &#8220;Stargate for Data&#8221;</strong> thread was the most substantive strategy post, arguing that data collection&#8212;not compute alone&#8212;becomes the binding constraint and potential moat for frontier labs <a href="https://x.com/willdepue/status/2074178395462848800">@willdepue</a>.</p></li><li><p><strong>John Carmack&#8217;s memory-system thread</strong> drew significant technical interest by arguing inference hardware could exploit deterministic access patterns and much cheaper memory tiers than HBM for large-model serving <a href="https://x.com/ID_AA_Carmack/status/2074248758422864226">@ID_AA_Carmack</a>.</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. Large Open-Weight MoE Model Releases</strong></h3><ul><li><p><strong><a href="https://www.reddit.com/r/LocalLLaMA/comments/1unyvnz/longcat_20_16t_48b_active_weights_are_now_open/">longcat 2.0 (1.6T, ~48B active) weights are now open under MIT license</a></strong> (Activity: 638): <strong>LongCat 2.0 weights are now open under the MIT license via announcements from <a href="https://x.com/eliebakouch/status/2073690402503487902">elie</a> and <a href="https://x.com/ModelScope2022/status/2073710226365165679">ModelScope</a>, with technical details in the <a href="https://longcat.chat/blog/longcat-2.0/">LongCat 2.0 blog post</a>. The model is a very large MoE system with </strong><code>1.6T</code><strong> total parameters and roughly </strong><code>48B</code><strong> active parameters per inference; commenters note the released weights occupy about </strong><code>3.55 TB</code><strong> in BF16 and </strong><code>2.05 TB</code><strong> in FP8.</strong> Commenters emphasized the practical deployment burden from the multi-terabyte weight size, and noted that <strong>Meituan</strong>&#8212;described as China&#8217;s Groupon/Uber Eats analogue&#8212;reportedly trained it on fully domestic Chinese chips, prompting discussion about the geopolitical/market significance.</p><ul><li><p>Commenters highlighted the scale and deployment footprint of <strong>LongCat 2.0</strong>: <code>1.6T</code> total parameters with approximately <code>48B</code> active parameters, implying a sparse/MoE-style architecture. One user noted the released weights require about <code>3.55 TB</code> in <strong>BF16</strong> and <code>2.05 TB</code> in <strong>FP8</strong>, which is important for anyone planning local storage or inference infrastructure.</p></li><li><p>A technical point raised was that <strong>Meituan</strong> reportedly trained the model on <code>100%</code> domestic Chinese chips, which commenters framed as significant for AI hardware supply-chain independence. This is especially notable given Meituan&#8217;s role as a major Chinese internet company comparable to a mix of Groupon and Uber Eats rather than a traditional AI lab.</p></li><li><p>Several users focused on the permissive <strong>MIT license</strong> and planned benchmarking against frontier open models such as <strong>Qwen</strong> and <strong>DeepSeek</strong>. The combination of <code>1.6T</code> total parameters, only <code>~48B</code> active parameters, and open weights suggests the model may be practical to compare with other high-end MoE open models if inference tooling supports its architecture efficiently.</p></li></ul></li></ul><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-the-field-guide-to-fable">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AIEWF Daily Dispatch: The great loops debate and the state of AI engineering]]></title><description><![CDATA[The AI Engineer World&#8217;s Fair ended with a debate about loops, a report on the state of AI engineering, and closing keynotes focused on what to build next.]]></description><link>https://www.latent.space/p/aiewf-daily-dispatch-locomotives</link><guid isPermaLink="false">https://www.latent.space/p/aiewf-daily-dispatch-locomotives</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Fri, 03 Jul 2026 05:11:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!M0WA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M0WA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M0WA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 424w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 848w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M0WA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg" width="1280" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:795551,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204783208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!M0WA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 424w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 848w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!M0WA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d4f7346-6688-4240-b077-16bf6f4a4a34_1280x815.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One of the highlights of the final day of the AI Engineer World&#8217;s Fair was a debate about loops. It nicely captured an argument running through the whole conference: are autonomous software factories viable now, or is the engineering discipline lagging behind the ambition?</p><p>Allie Howe from Keycard was the moderator and she opened by asking, &#8220;is there or is there not a delta between the hype behind loops and what actually works in practice?&#8221;</p><p>The pro-loop case was presented by Geoffrey Huntley, creator of the <a href="https://ghuntley.com/loop/">Ralph Loop</a>, and Keycard CEO Ian Livingstone. Huntley opened by saying loops are already here. &#8220;It&#8217;s inevitable, it&#8217;s here to stay,&#8221; adding that &#8220;I don&#8217;t see myself going back to writing code by hand.&#8221;</p><p>Livingstone said that verifiability is ultimately what it&#8217;s about &#8212; and you can achieve that with any code, regardless of how it was produced. He also pointed out that loops have always been a core aspect of software development:</p><p>&#8220;A loop is at the core of &#8216;I try something, I learn something, I apply something.&#8217; And all we&#8217;re really talking about is how quickly we can expedite that process.&#8221;</p><p>On the skeptical side were Dex Horthy from HumanLayer and Greg Pstrucha from Subroutine. Horthy began by noting that he wasn&#8217;t anti-loops. &#8220;The basic take here is not whether loops are good or bad,&#8221; he said, noting that &#8220;Kubernetes is actually built on loops &#8212; built on control loops. But they&#8217;re deterministic loops.&#8221; Horthy&#8217;s issue is that &#8220;the hype is outrunning the discipline.&#8221;</p><p>&#8220;I haven&#8217;t seen proof that we are at a point where we can just step up an abstraction level,&#8221; Horthy said, referring to agents controlling the coding. &#8220;I actually think we need to step down an abstraction level, if anything.&#8221;</p><p>Pstrucha was mainly concerned about the economic viability of agentic loops, which he said wasn&#8217;t sustainable. You can&#8217;t &#8220;orchestrate your problems away by buying more tokens,&#8221; he said.</p><div class="pullquote"><p>&#8220;[We&#8217;re] kind of like locomotive engineers now. That&#8217;s our job: to keep the locomotive on the rails.&#8221;<br>- Geoffrey Huntley, loops advocate</p></div><p>Huntley then offered this wonderful analogy for loopmaxxing: &#8220;[We&#8217;re] kind of like locomotive engineers now. That&#8217;s our job: to keep the locomotive on the rails.&#8221;</p><p>The discussion turned to <a href="https://www.latent.space/p/software-factories">software factories</a>, the metaphor that has really taken hold of the industry. Horthy worries that when everything is automated in a factory-like agent environment, &#8220;you never touch the problem.&#8221; So instead, he advises to start small and iterate with agent loops &#8212; to &#8220;build up intuition&#8221; and not try to automate end to end from the start.</p><p>Even Huntley recognized some of the dangers in loops. He said that software factories represent where we are headed in the future, but cautioned that it&#8217;s not yet solved in the market. &#8220;This is frontier thinking,&#8221; he said.</p><p>At the end of the hour-long debate, Howe polled the audience to ask which side &#8216;won&#8217;. Ironically, this resulted in a human failure: the stage lights were too bright for Howe or any of the debate participants to see how many hands were raised. If only an agent was in charge of dimming the lights.</p><h2>Anthropic&#8217;s next big thing: Claude Tag</h2><p>Perhaps one example of a company moving to a software factory model is Anthropic. Mike Krieger, one of the co-founders of Instagram back in Web 2.0 and now Head of Labs at Anthropic, was interviewed by swyx in one of the morning sessions.</p><p>Krieger talked about <a href="https://www.anthropic.com/news/introducing-claude-tag">Claude Tag</a>, Anthropic&#8217;s internal model which the company announced to the world last week. He described Tag as more delegated, asynchronous and proactive than Claude. It perhaps suggests what an early software factory looks like in practice &#8212; not agents replacing a team, but multiple people delegating responsibilities to a system like Claude Tag.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Av0r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Av0r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Av0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg" width="1280" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:880968,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204783208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Av0r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Av0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26adc828-7702-490a-be57-f91ac1adb699_1280x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Mike Krieger talking with swyx at AIEWF today.</figcaption></figure></div><p>&#8220;Most usage is actually much more delegated,&#8221; he said regarding his team&#8217;s usage of Tag. He gave an example of how they instruct the agents: &#8220;Don&#8217;t just fix this bug. Now you are responsible for this part of the codebase, and I want you to monitor this feedback channel and proactively take on tasks.&#8221;</p><p>&#8220;That&#8217;s really changed how we operate currently,&#8221; he continued. &#8220;It&#8217;s much more this multiplayer, async, proactive way.&#8221;</p><p>However, he also indicated there are some negative consequences to becoming more automated. He noted that his team is &#8220;bottlenecked on reviews&#8221; and on the &#8220;human ability to fully conceptualize what we&#8217;re doing.&#8221;</p><h2>2026 AI Engineer Survey</h2><p>Back to the current reality for most AI engineers. This morning, Barr Yaron from Amplify presented her annual survey of the industry.</p><p>According to Amplify&#8217;s data, 95% of respondents now use agents &#8212; roughly double last year&#8217;s share. Among teams using agents, 89% said those agents could write data, up from 52% the previous year.</p><p>&#8220;Agents are no longer reading, summarizing, drafting,&#8221; Yaron said. &#8220;They&#8217;re taking actions inside the systems.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zAbI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zAbI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zAbI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg" width="1280" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:744512,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204783208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zAbI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!zAbI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7fe746-5558-4484-92a2-82f5309f0230_1280x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Barr Yaron presenting her AI engineering survey.</figcaption></figure></div><p>The controls, however, remain comparatively primitive. Human approvals and permissions were the two leading safeguards, followed by a scattered collection of task decomposition, retrieval, memory and sandboxing techniques.</p><p>&#8220;Nobody has settled the control layer for agents,&#8221; Yaron said.</p><p>Cost is also a concern. Forty percent of respondents said that AI costs regularly limit how ambitiously they use AI, while another 36% said it sometimes does. Token usage is now the second-most monitored production metric, behind quality.</p><p>The survey captured the conference&#8217;s central contradiction. AI has made experimentation cheaper and enabled teams to produce more software, but 59% of respondents to the Amplify survey fear that today&#8217;s AI-generated code is creating long-term liabilities.</p><h2>Closing keynotes</h2><p>The final sessions of the conference appropriately took us back to thinking optimistically about AI technology &#8212; about building with it. After all, that&#8217;s why the AI Engineer World&#8217;s Fair exists, and it&#8217;s where the fun is! </p><p>Theo Browne showcased several software projects he had built, or was still building, with AI. His point was that the scale of what an individual developer can realistically attempt has shifted. &#8220;What used to be a startup is now a side project,&#8221; he said, while projects he would once have dismissed as &#8220;too big&#8221; are moving within reach. </p><p>Garry Tan, president and CEO of Y Combinator, followed by giving that optimism an organizational form. The fastest-growing founders YC sees, he said, are &#8220;not treating AI as autocomplete, they&#8217;re treating it as a workforce.&#8221; </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TVDr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TVDr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TVDr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg" width="1280" height="750" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:750,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:508140,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204783208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TVDr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TVDr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f4f4636-6d5b-422a-97c9-e653e4f100e3_1280x750.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Garry Tan at AIEWF.</figcaption></figure></div><p>Tan&#8217;s closing prescription was: &#8220;Build an AI-native company, not a company that just uses AI.&#8221;</p><p>The debates during the week showed how much engineering remains before the AI-native vision is viable for all. But the closing keynotes offered a reminder of why the engineers who attended this conference are pursuing it: they just want to ride those locomotives!</p>]]></content:encoded></item><item><title><![CDATA[Vercel's Andrew Qu on why agents are a new kind of software]]></title><description><![CDATA[The Vercel Chief of Software explains how its agent framework, eve, was created &#8212; and why skills, sandboxes and agent-readable websites now matter.]]></description><link>https://www.latent.space/p/vercel-agents-new-software</link><guid isPermaLink="false">https://www.latent.space/p/vercel-agents-new-software</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Fri, 03 Jul 2026 00:08:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5d05f30e-e2cc-4895-b84a-d0cdd9835db8_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-oUP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-oUP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-oUP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:893246,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204762364?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-oUP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-oUP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faba6a476-7db8-4211-967c-a4f1ca928e66_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Vercel&#8217;s Andrew Qu on the AIEWF expo floor.</figcaption></figure></div><p><a href="https://x.com/andrewqu">Andrew Qu</a> is Chief of Software at Vercel, where he works with the CTO across internal engineering, product experimentation and emerging technologies. He has built libraries for MCP, created skills.sh and led the development of eve, Vercel&#8217;s framework for building agents.</p><p>In this interview with Latent Space, Qu explains why agents represent a new form of software, what Vercel learned from building its own, and why Vercel itself is turning into an agent!</p><h2>From web applications to agents</h2><p><strong>Latent Space: </strong>What does a Chief of Software do at Vercel?</p><p><strong>Andrew Qu:</strong> My role is pretty unique. I work with the CTO to ship impact in any way, shape or form. It&#8217;s a mix of internal engineering, external experimentation and staying on the frontier by building things.</p><p>That means building new libraries and frameworks and showing people how to do things for the first time. I built an MCP library that made it easier to create some of the first MCP servers, and I also built skills.sh to make agent skills easier to discover and use.</p><p><strong>Latent Space: </strong>How did Vercel evolve from focusing on web development to investing heavily in agents?</p><p><strong>Qu:</strong> Vercel&#8217;s origins were about making it easy for developers to ship websites and web applications. More recently, we&#8217;ve seen a shift from people building pages to people building agents.</p><p>While building our own agent in v0, our vibe-coding product, we ran into a lot of paper cuts that existing tooling did not solve: switching models or providers, adding fallbacks and making runs resumable.</p><p>We turned those solutions into reusable libraries that could support v0 and also help customers build their own agents. Over time, we accumulated a set of primitives and decided to assemble them more cohesively. That became eve.</p><h2>Why eve became necessary</h2><p><strong>Latent Space: </strong>How did you reach the point where Vercel needed a dedicated agent framework?</p><p><strong>Qu:</strong> About a year ago, I started working toward putting an agent on every desk inside Vercel. That led me to build a successful data agent, and along the way a number of best practices emerged: filesystem agents, skills, compaction and subagents.</p><p>These were all things I wished had come out of the box. Eventually, we asked: what if there were a prescriptive way to do this, so other developers did not have to go through the same exploration? That is where eve came from.</p><p><strong>Latent Space: </strong>Are agents simply another kind of application, or a genuinely new form of software?</p><p><strong>Qu:</strong> I think agents are a new type of software. They are not as predictable as web applications. The infrastructure can look similar, but the interaction, interface and outputs are much more dynamic.</p><p>That changes how you build them. You need different primitives for context, tools, resumability and long-running work.</p><p><strong>Latent Space: </strong>What kinds of problems are particularly well suited to agents?</p><p><strong>Qu:</strong> We see a lot of business agents. Internally at Vercel, we use them for repetitive work ranging from a first pass at legal contract redlining, to marketing retrospectives and identifying people to contact, to writing queries against our data stores.</p><p>A good candidate is often a repetitive task that still requires some reasoning. It is not just fixed automation, because the system has to interpret the situation and decide what to do.</p><h2>Building effective agents</h2><p><strong>Latent Space: </strong>When should an agent work autonomously, and when should a human remain in the loop?</p><p><strong>Qu:</strong> I don&#8217;t think the future is all autonomous loops, and I don&#8217;t think it is all human-in-the-loop. It is about choosing a feedback cycle that fits the task.</p><p>If the task is well defined and you know what the final output should look like, it can be reasonable to let a loop continue until it is done. For more careful or surgical engineering work, you should check back in and make sure you are steering the model correctly.</p><p><strong>Latent Space: </strong>Your approach evolved through prompting, bespoke tools, coding-agent harnesses, filesystem agents and skills. What was the main lesson?</p><p><strong>Qu:</strong> We are still figuring out what makes an agent productive. Along the way, we have been collecting these primitives and bringing them together in eve.</p><p>There will be more to add as best practices emerge. A year ago, we did not know sandboxes would become so important, or how much demand there would be for secure code execution and long-running jobs. As we learn more from production, there will be much more to build.</p><p><strong>Latent Space: </strong>Is Vercel creating an end-to-end agent platform comparable to the one it built for web development?</p><p><strong>Qu:</strong> Yes and no. We value partners that provide specialized parts of the agent lifecycle, but we also want it to be very easy for developers to get started.</p><p>If you deploy eve to Vercel, you get observability and evaluations out of the box. We want to make that experience more comprehensive while making it easy to integrate with partners rather than owning every component.</p><h2>Skills and current knowledge</h2><p><strong>Latent Space: </strong>Why have skills become so important?</p><p><strong>Qu:</strong> Skills are useful as portable, on-demand knowledge. Models often contain outdated information. For example, they still sometimes recommend Vercel Postgres, even though we deprecated it years ago in favor of our marketplace.</p><p>A skill can tell the agent that Vercel Postgres is deprecated and steer it toward the current approach. Until companies can audit and update every old piece of content, skills provide a way to forward-correct the model.</p><p>I would recommend publishing skills for the latest version of your product. But companies should also audit their existing content, identify what is outdated and update it or add clear notes.</p><h2>An agent-readable web</h2><p><strong>Latent Space: </strong>How will websites evolve as more traffic comes from agents?</p><p><strong>Qu:</strong> We have published reports showing bot traffic rising while human traffic is stagnant or declining, even as impressions increase, because agents and bots are hitting websites more frequently.</p><p>The future of the web is therefore to be as accessible to bots and agents as possible, so they can learn about your product and use it successfully.</p><p>At Vercel, we already detect when an agent makes a request and serve Markdown directly. Instead of forcing it to process HTML designed for a visual browser, we provide a format that is easier to read.</p><p><strong>Latent Space: </strong>Does that mean one experience for humans and another for agents?</p><p><strong>Qu:</strong> I think so. Humans may continue to receive the visual site, while agents receive a more structured, machine-readable representation. We are already doing that today.</p><h2>What comes next</h2><p><strong>Latent Space: </strong>What problems are you most interested in solving next?</p><p><strong>Qu:</strong> One of the things at the top of my agenda is multiplayer agent development. Whenever a team collaborates, people struggle to share context.</p><p>I may have techniques for getting a front-end interface right on the first attempt, but another person may not know them. I am interested in how we can share that context between teammates and allow them to contribute to it.</p><p><strong>Latent Space: </strong>Will agents become a separate application category, or a standard capability built into most software?</p><p><strong>Qu:</strong> It depends on who you are and what you are building. For Vercel, Vercel itself is becoming an agent. We have an agent on the website, in Slack and in the dashboard that can do things on your behalf.</p><p>Other companies will ship agents as standalone products. For us, agents are tightly coupled to everything we build. We want the entire platform to be agent-friendly &#8212; and, in many ways, to make the platform itself an agent.</p>]]></content:encoded></item><item><title><![CDATA[The website of the future may assemble itself for every visitor]]></title><description><![CDATA[Adobe is experimenting with &#8220;agentic sites&#8221; that generate pages around an individual user&#8217;s intent. At AIEWF, we talked to Carlos Sanchez about the Web's future.]]></description><link>https://www.latent.space/p/the-website-of-the-future</link><guid isPermaLink="false">https://www.latent.space/p/the-website-of-the-future</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Thu, 02 Jul 2026 21:25:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KiR3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KiR3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KiR3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KiR3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:646204,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204745876?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KiR3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KiR3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F888e1825-8152-418b-9a18-152de1328077_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Adobe Principal Scientist Carlos Sanchez at AIEWF.</figcaption></figure></div><p>For as long as I can remember (and I managed websites in the dot-com period), &#8220;personalization&#8221; has been a holy grail for websites. But up till now, that&#8217;s typically meant selecting from a predefined set of options. A retailer might recommend an item based on a previous purchase, or place a visitor into one of several audience segments &#8212; that&#8217;s been the extent of personalization.</p><p>Adobe Principal Scientist <a href="https://x.com/csanchez">Carlos Sanchez</a> is exploring a more radical possibility: what if the website itself could be assembled around the needs of each visitor?</p><p>At the AI Engineer World&#8217;s Fair in San Francisco, Sanchez demonstrated what Adobe calls an &#8220;agentic site&#8221; &#8212; a web experience that interprets a visitor&#8217;s intent, retrieves relevant material from the company&#8217;s existing content, and composes a personalized page in real time.</p><p>Adobe calls this approach an &#8220;audience of one.&#8221; Sanchez&#8217;s larger point was that the technology is no longer hypothetical.</p><p>&#8220;Many people don&#8217;t even think it&#8217;s possible to generate a web page on the fly,&#8221; he told Latent Space after his session. &#8220;People think it is future-looking. No, you can do this. It&#8217;s not the future, it&#8217;s the present now.&#8221;</p><h2><strong>From personalized components to personalized pages</strong></h2><p>During his presentation, Sanchez demonstrated a site that used the visitor&#8217;s browsing behavior and search queries as signals. The system grouped those signals into an intent category &#8212; such as exploring, researching or preparing to purchase &#8212; and then used an LLM to assemble a page suited to that intent.</p><p>In one example, a visitor interested in camping received a version of a coffee-machine site whose copy, product selection and supporting content had been reorganized around making coffee outdoors.</p><p>Sanchez also showed a more open-ended interface in which someone could enter a query such as &#8220;Europe AI conferences&#8221; and receive a page composed specifically around that request.</p><p>&#8220;We call this &#8216;audience of one,&#8217; because the idea is to personalize the site in real time based on the user accessing it and what the user is doing,&#8221; Sanchez said.</p><p>The idea is that the site&#8217;s existing content is the grounding corpus. Adobe&#8217;s system retrieves from that material rather than asking an LLM model to invent an entire experience from scratch.</p><p>For AI engineers, one potential constraint is latency. In his session, Sanchez said that Adobe evaluates models not only for accuracy, but also for speed: &#8220;We don&#8217;t want the site generation to take more than one or two seconds.&#8221;</p><p>Sanchez says the economics are already becoming plausible. He estimated the current inference cost at &#8220;one to two cents per page.&#8221;</p><p>&#8220;But our point is also this is only going to get cheaper,&#8221; he said. &#8220;This is where we are today. In six months, who knows where we&#8217;re going to be.&#8221;</p><h2><strong>AI makes it easier to build, but harder to choose</strong></h2><p>Adobe has not yet broadly deployed these experiences on production customer sites. Sanchez said the company is presenting the concept to customers and looking for organizations willing to experiment.</p><p>Commerce is an obvious initial use case, because personalization can be connected directly to conversion. But the opportunity is not necessarily limited to retail. &#8220;It could work for other things &#8212; anything that needs more conversion and has a big matrix of user types or personas,&#8221; he told me.</p><p>Still, Sanchez acknowledged that he&#8217;s unsure if agentic sites will become a widespread reality.</p><p>&#8220;With AI, it&#8217;s very easy to build things, but it&#8217;s hard to know what to build,&#8221; he said. &#8220;We build things and then we find the customers.&#8221;</p><p>It&#8217;s not just Adobe feeling the uncertainty around its &#8216;audience of one&#8217; concept. Website owners are currently evaluating all kinds of AI functionality: chat interfaces, structured content (like WebMCP), generative UI, personal agents, and more. Not to mention trying to find ways to bring users in from third-party AI platforms.</p><p>&#8220;I think it&#8217;s a combination of all these crazy different ways,&#8221; Sanchez said. &#8220;You are in a chat, I want to show UI, I want to get you to buy something. Then you&#8217;re in a site, I want to steer you this other way. Maybe you&#8217;re in an OpenAI chat and I want to bring you into my site. Everybody&#8217;s trying to figure this out on the marketing side.&#8221;</p><h2><strong>A web built for humans &#8212; and agents</strong></h2><p>Of course, websites in 2026 and beyond won&#8217;t just be personalized for human visitors.</p><p>As personal agents become more capable, a user may delegate some purchases or research tasks entirely. The agent could arrive carrying a much richer expression of the user&#8217;s preferences than the destination site could infer from cookies or recent browsing behavior.</p><p>Sanchez expects websites to evolve for both kinds of visitor. &#8220;Whether it&#8217;s going to be two versions [of a website] or not, that may be blurry,&#8221; he said. &#8220;But obviously, you&#8217;re going to have to target both.&#8221;</p><p>Also, not every transaction will work the same way. A personal agent might autonomously reorder toilet paper, while a person buying a jacket may still want to inspect the product and make the final choice through a visual interface.</p><p>That means websites will need to support different levels of delegation and involvement, rather than treating &#8220;agentic commerce&#8221; as a single interaction pattern.</p><p>Technologies such as WebMCP could allow a site to expose structured tools directly to an agent, while MCP Apps and other generative interfaces could bring interactive product experiences into the user&#8217;s chat environment. An A2A backend might allow agents to interact without traversing the conventional visual site at all.</p><p>It might end up being one site with both visual components and agent-accessible tools &#8212; two distinct experiences &#8212; or perhaps a human-facing website paired with an agent-to-agent service.</p><p>&#8220;That&#8217;s still what everybody&#8217;s trying to figure out,&#8221; Sanchez said. &#8220;But there&#8217;s going to be agentic targeting, for sure.&#8221;</p><h2>Whither websites?</h2><p>Whether websites survive the AI era at all is another big question we&#8217;re all grappling with.</p><p>What I gleaned from Sanchez at AIEWF was that the traditional website is unlikely to disappear completely, but its role will surely change.</p><p>Rather than being a fixed collection of pages that every visitor navigates, a &#8220;website&#8221; could become a governed content and interaction system that assembles an appropriate interface on demand. At least, that&#8217;s the future that Adobe is actively exploring.</p>]]></content:encoded></item><item><title><![CDATA[Skill engineering and the case against one-shot AI design]]></title><description><![CDATA[Paul Bakaus talks to us about Impeccable, human judgment in a 'loopmaxxing' era, and why agents still need people to steer them.]]></description><link>https://www.latent.space/p/skill-engineering-design</link><guid isPermaLink="false">https://www.latent.space/p/skill-engineering-design</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Thu, 02 Jul 2026 14:36:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JOvz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JOvz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JOvz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JOvz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:682828,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204688240?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JOvz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JOvz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c23df3d-275c-48a8-a914-994c53dcd352_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Impeccable&#8217;s Paul Bakaus at the AI Engineer World&#8217;s Fair.</figcaption></figure></div><p>Paul Bakaus thinks the emerging discipline of &#8220;skill engineering&#8221; can make AI agents more capable &#8212; but he absolutely does <strong>not</strong> want to remove people from the creative process. He chats to Latent Space about his approach to design in the AI age.</p><p><a href="https://www.paulbakaus.com/">Bakaus</a> is the creator of <a href="https://impeccable.style/">Impeccable</a>, an open-source design skills system that gives coding agents a vocabulary for improving interfaces. Instead of asking an agent to redesign an entire website in one shot, users can tell it to make a section &#8220;bolder,&#8221; &#8220;quieter,&#8221; &#8220;denser,&#8221; or more polished.</p><p>Behind those apparently simple commands is a larger argument about how AI products should be built. Agents need more than instructions, Bakaus said: they need domain knowledge, context and carefully defined ways for humans to steer the result.</p><p>&#8220;The point is to give you a way to steer what you want to end up with,&#8221; he said during a session at the AI Engineer World&#8217;s Fair. &#8220;It&#8217;s never going to be a tool for one-shot design. That&#8217;s not the intent.&#8221;</p><h2><strong>The emerging craft of skill engineering</strong></h2><p>Impeccable began as a relatively simple extension of Anthropic&#8217;s frontend design skill. As its audience grew, Bakaus expanded it into a more complex system with multiple components and workflows.</p><p>That process led him to start thinking of skill engineering as a discipline in its own right. His workshop at the conference explored what he called the &#8220;dark arts&#8221; of building skills.</p><p>&#8220;One of the interesting topics was that most skills &#8212; [and] most models &#8212; are not very creative,&#8221; Bakaus told me. &#8220;They converge in one direction, and if everybody uses the same skill to do frontend design work or something like that, everything ends up looking the same.&#8221;</p><p>Skill engineers must also account for differences between agent harnesses and models. Codex and Claude, for example, do not necessarily handle subagents or permissions in the same way. A skill intended to run across Claude Code, Cursor, GitHub Copilot and Codex cannot assume they all provide identical capabilities.</p><p>Bakaus has also experimented with routing inside a skill, allowing it to combine several capabilities and direct a task toward the relevant instructions. He compared this to a mixture-of-experts model, with routing used both to conserve tokens and improve effectiveness.</p><h2><strong>Giving agents a design vocabulary</strong></h2><p>Impeccable&#8217;s core innovation is to take terms familiar to designers and give them a more precise operational meaning for an agent.</p><p>An unassisted model asked to make a page &#8220;bolder&#8221; may add gradients, neon effects or glass-like surfaces. Impeccable instead defines boldness through concepts such as hierarchy, scale and decisive typography &#8212; changes that attract attention without necessarily breaking the existing design system.</p><p>&#8220;An adjective with nothing behind it is just a nice apostrophe,&#8221; Bakaus said. &#8220;You really have to tell the agent what you mean.&#8221;</p><p>He described these terms as words that have been &#8220;imbued with meaning.&#8221; The model already has some conception of what words such as &#8220;bold&#8221; or &#8220;quiet&#8221; mean, but the skill translates them into a specific professional domain.</p><p>This is the key, because experts often possess a vocabulary that non-experts do not. Bakaus said he had observed large differences between the work produced by a designer and an engineer using the same model, simply because the designer knew how to articulate the desired result.</p><p>&#8220;I&#8217;ve been trying to put that language &#8212; basically compress it into a skill and into a system &#8212; to be able to express yourselves better,&#8221; he said.</p><p>However, he does not believe every part of design can be controlled from this level of abstraction. Directly manipulating spacing may still be the fastest option for a small adjustment, while open-ended prompting can be useful during initial exploration.</p><p>The objective is not to replace every tool with an agent, he insisted. It is to determine &#8220;the exact level of control&#8221; and insert the person at the point where their judgment is most valuable.</p><h2><strong>Designers and engineers move up the stack</strong></h2><p>Bakaus sees the boundaries between design, engineering and product management becoming less distinct.</p><p>&#8220;Designers are moving into code, engineers are moving into design, and vice versa,&#8221; he said. &#8220;These worlds are all colliding.&#8221;</p><p>That shift will be uncomfortable for people whose work primarily consists of translating an existing artifact into another form. Engineers who mainly turn Figma designs into code face growing automation, while designers whose contribution is limited to making an existing interface look competent face similar pressure.</p><p>&#8220;Designers all have to move one layer up the stack to think more about the <em>what</em>,&#8221; he said. &#8220;I think the role of the product manager and designer is actually converging.&#8221;</p><p>At the same time, designers are moving closer to implementation &#8212; into code. Bakaus initially expected Impeccable to appeal mostly to engineers and assumed professional designers might resent that. Instead, he estimates that designers now make up at least half of its audience.</p><p>&#8220;So rather than moving directly into code and, you know, having no help,&#8221; Bakaus said about designers, &#8220;they use Impeccable as a bridge, because it communicates the way they communicate. And that was not obvious to me when I first built it.&#8221;</p><p>Impeccable also has a live mode that combines visual selection with an underlying coding agent. A user can select a section inside a development environment and request several alternative layouts or (for example) ask for a bolder or quieter treatment. The system operates within the project&#8217;s existing code and design system rather than exporting an isolated mockup from a third-party design tool.</p><p>Bakaus described this as a potential &#8220;design harness&#8221; at the intersection of chat and direct visual manipulation.</p><h2><strong>There will be no auto mode</strong></h2><p>The AI industry often treats complete automation as the natural endpoint of product development. Bakaus rejects that premise.</p><p>He sees two dominant camps: people trying to preserve the traditional Figma-centered workflow, and on the other side advocates of &#8220;loopmaxxing&#8221; who want agents to work with as little human intervention as possible.</p><p>&#8220;The truth is somewhere in the middle,&#8221; he said.</p><p>His preferred model is for AI to produce the first 80% quickly: the competent layout and basic implementation that would otherwise consume a lot of time. The person then owns the final 20%, where taste, context and a distinctive point of view enter the product. This is a key part of Bakaus&#8217;s design philosophy in the agentic era.</p><p>&#8220;People need purpose, and they want to play a role in whatever they create,&#8221; Bakaus said. &#8220;When you work with the agent, then you feel more ownership of the product.&#8221;</p><p>Users regularly ask him to add an automatic mode to Impeccable so that the system chooses the commands itself. He has no intention of doing so.</p><p>&#8220;There is no auto,&#8221; he said, &#8220;and there will be no auto.&#8221;</p><p>Asked about the language of <a href="https://www.latent.space/p/software-factories">software factories</a> and other visions that appear to remove people from engineering altogether, his response was unambiguous.</p><p>&#8220;I&#8217;m squarely against that.&#8221;</p>]]></content:encoded></item><item><title><![CDATA[[AINews] not much happened today]]></title><description><![CDATA[another quiet day.]]></description><link>https://www.latent.space/p/ainews-not-much-happened-today-900</link><guid isPermaLink="false">https://www.latent.space/p/ainews-not-much-happened-today-900</guid><pubDate>Thu, 02 Jul 2026 07:10:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/4sX_He5c4sI" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Fable was relaunched on schedule, and AIE was on top of it with <strong>the first Field Guide to Fable talk</strong>, as well as the rest of the excellent <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Richard MacManus&quot;,&quot;id&quot;:232063,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c4ca3255-4ccf-497e-a04f-219d65fba554_2048x2048.png&quot;,&quot;uuid&quot;:&quot;3b48ee57-ebcc-47a5-9873-18637cecec64&quot;}" data-component-name="MentionToDOM"></span> coverage of AIEWF Day 3 across <a href="https://www.latent.space/p/autoresearch-introspection">Autoresearch</a>, <a href="https://www.latent.space/p/cursor-forward-deployed-engineers">Cursor FDE</a>, and a <a href="https://www.latent.space/p/software-factories">followup</a> to <a href="https://www.youtube.com/watch?v=4sX_He5c4sI">Zach Lloyd&#8217;s popular talk yesterday on Software Factories</a>, as well as &#8220;all killer no filler&#8221; closing keynotes:</p><div id="youtube2-4sX_He5c4sI" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;4sX_He5c4sI&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/4sX_He5c4sI?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><blockquote><p>AI News for 7/1/2026-7/1/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Coding Models, Agent Harnesses, and the Fable 5 Re-launch</strong></p><ul><li><p><strong>Anthropic re-enabled Claude Fable 5, but with visible safety fallbacks</strong>: After a day of pent-up demand, <a href="https://x.com/claudeai/status/2072402636813607381">@claudeai</a> announced <strong>Fable 5 is back</strong>, alongside a clarifying note that updated cybersecurity safeguards may route some requests to <strong>Opus 4.8</strong>, with biology/chemistry classifiers still overly broad for now <a href="https://x.com/claudeai/status/2072402638247968855">@claudeai</a>. The relaunch immediately propagated into tooling: <strong>Cursor</strong> says Fable 5 leads its evals but is the <strong>most expensive per task</strong> <a href="https://x.com/cursor_ai/status/2072403323844428217">@cursor_ai</a>; <strong>Devin</strong> added it across Cloud/Desktop/CLI <a href="https://x.com/cognition/status/2072405137117548601">@cognition</a>; <strong>Perplexity</strong> restored it as an orchestrator model <a href="https://x.com/perplexity_ai/status/2072433125104505226">@perplexity_ai</a>. Anthropic also reset rate limits for users once the model was live again <a href="https://x.com/ClaudeDevs/status/2072429181565288665">@ClaudeDevs</a>.</p></li><li><p><strong>The interesting story was less &#8220;model is back&#8221; than &#8220;how people are adapting to frontier-model constraints&#8221;</strong>: Multiple builders converged on <strong>multi-model orchestration</strong> rather than single-model dependence. <a href="https://x.com/theo/status/2072481845363822914">@theo</a> described using Fable only for higher-value reasoning/planning while delegating implementation, verification, and computer-use work to other models; he reports a substantial improvement in end-to-end PR yield <a href="https://x.com/theo/status/2072482460122964067">@theo</a>. Similar views came from <a href="https://x.com/omarsar0/status/2072400978079261041">@omarsar0</a>, who argued teams should design <strong>model-combination strategies</strong> rather than build around one frontier model, and from <a href="https://x.com/MParakhin/status/2072275413116784961">@MParakhin</a>, who pushed back on &#8220;simple-task pre-classifiers,&#8221; arguing that reliable routing often requires solving the task first. On the benchmark side, <a href="https://x.com/kimmonismus/status/2072376968729817531">@kimmonismus</a> highlighted <strong>Fable 5&#8217;s 16.10% on the Remote Labor Index</strong>, while <a href="https://x.com/ArtificialAnlys/status/2072427328689619241">@ArtificialAnlys</a> reported <strong>Sonnet 5</strong> ranking second on <strong>AA-Briefcase</strong> but with much higher turn counts and weaker cost-performance tradeoffs at lower effort settings.</p></li></ul><p><strong>Open Models, Chinese Labs, and the Expanding Coding Stack Around GLM-5.2</strong></p><ul><li><p><strong>Z.ai is building product surface area around GLM-5.2, not just shipping a checkpoint</strong>: The most concrete launch was <strong>ZCode</strong>, the official dev environment for <strong>GLM-5.2</strong>, with BYOK support, cross-platform availability, and a quota boost for coding-plan subscribers <a href="https://x.com/Zai_org/status/2072349453361557898">@Zai_org</a>. Commentary from <a href="https://x.com/kimmonismus/status/2072378141041991702">@kimmonismus</a> framed it as an AI-native coding IDE optimized for GLM workflows and long-running autonomous tasks. The surrounding ecosystem is moving quickly too: <strong>LangChain</strong> published guides for using GLM-5.2 in coding flows <a href="https://x.com/LangChain/status/2072334663457067064">@LangChain</a>, and <a href="https://x.com/hwchase17/status/2072344890755977571">@hwchase17</a> explicitly called out developers turning to GLM-5.2 as a daily driver.</p></li><li><p><strong>Benchmarks suggest open coding models are closing specific gaps even if not leading overall frontier performance</strong>: <a href="https://x.com/mercor_ai/status/2072448918751941041">@mercor_ai</a> reported <strong>GLM 5.2</strong> as the first open model to lead a category on <strong>APEX-SWE</strong>, posting <strong>55.3% Pass@1 on Integration</strong>, and ranking as the best open model tested overall there; <strong>Kimi K2.7</strong> followed closely. That complements <a href="https://x.com/scaling01/status/2072346101068238946">@scaling01</a>, who cautioned against overclaiming that GLM has surpassed top Western frontier models while still acknowledging a rapidly shrinking coding gap.</p></li><li><p><strong>Inference work around open models is becoming a meaningful part of the story</strong>: <a href="https://x.com/vllm_project/status/2072545387639189798">@vllm_project</a> landed native <strong>DSpark speculative decoding</strong> support in <strong>vLLM</strong> for DeepSeek models, reporting around <strong>250 tok/s</strong> on 8&#215;B300 with improved acceptance over MTP, and <a href="https://x.com/mgoin_/status/2072525522639212825">@mgoin_</a> released a <strong>GLM-5.2 DSpark preview</strong> claiming roughly <strong>1.5&#215; faster decode</strong>. Separately, <a href="https://x.com/jon_durbin/status/2072293557172363720">@jon_durbin</a> reported an in-house <strong>dflash</strong> drafter on <strong>Qwen3-32B</strong> yielding <strong>~50% higher throughput</strong> on the same hardware.</p></li></ul><p><strong>Agent Infrastructure: Memory, Wikis, Skill Composition, and Structured Workflows</strong></p><ul><li><p><strong>&#8220;Wiki memory&#8221; is emerging as a practical design pattern for agents</strong>: <a href="https://x.com/sydneyrunkle/status/2072311589072486879">@sydneyrunkle</a> argued for <strong>wiki-structured memory</strong> as a simple, extensible substrate, and that idea rapidly turned into product releases. <strong>LangChain</strong> launched <strong>OpenWiki</strong>, a tool to generate and maintain agent-consumable codebase docs with <code>openwiki --init</code> <a href="https://x.com/BraceSproul/status/2072375499125596262">@BraceSproul</a>, <a href="https://x.com/LangChain/status/2072376975545798792">@LangChain</a>. The motivation is consistent across posts: agents repeatedly lose working context between threads and need a maintained, inspectable knowledge layer rather than raw logs <a href="https://x.com/caspar_br/status/2072420582717858292">@caspar_br</a>.</p></li><li><p><strong>Memory systems are shifting from retrieval-only to reconciliation and maintenance</strong>: Weaviate&#8217;s <strong>Engram</strong> pitch is representative here: candidate memories are extracted, transformed against existing memory, and only then committed, so contradictions are resolved once rather than at every query <a href="https://x.com/PrajjwalYd/status/2072291317695324410">@PrajjwalYd</a>. <a href="https://x.com/bpalit/status/2072378273343082537">@bpalit</a> extends the same argument to enterprise settings, where agent memory must be governed, permission-aware, and shared&#8212;not just a folder of markdown files.</p></li><li><p><strong>Structured composition is replacing naive &#8220;give the model all the tools&#8221; approaches</strong>: <a href="https://x.com/omarsar0/status/2072430551446032847">@omarsar0</a> highlighted <strong>SkillComposer</strong>, which treats skill selection as a joint autoregressive composition problem and reports <strong>+23.1pp / +18.2pp</strong> gains on SkillsBench over no-skill baselines. On the framework side, Deep Agents added support for <strong>recursive language model workflows</strong> <a href="https://x.com/sydneyrunkle/status/2072348322526810594">@sydneyrunkle</a>, and <a href="https://x.com/hwchase17/status/2072377816780624266">@hwchase17</a> connected <strong>dynamic subagents</strong> to patterns like <strong>Agentic MapReduce</strong>. This general direction&#8212;more explicit workflow structure, fan-out/fan-in patterns, and code-enforced orchestration&#8212;showed up repeatedly across products and benchmarks.</p></li></ul><p><strong>Security, Evaluation, and Agentic MapReduce</strong></p><ul><li><p><strong>Cognition&#8217;s Devin Security Swarm is one of the clearer examples of agent architecture specializing around a real enterprise workflow</strong>: The system uses <strong>Agentic MapReduce</strong> to fan out bounded agents across a codebase, aggregate findings, and validate exploitability before surfacing confirmed vulnerabilities <a href="https://x.com/cognition/status/2072368168182432109">@cognition</a>. Cognition claims this is both <strong>more cost-effective and more accurate</strong> than alternatives, and says a Fortune 500 pilot found and fixed <strong>over a thousand vulnerabilities</strong> in production repos <a href="https://x.com/walden_yan/status/2072377406267273248">@walden_yan</a>. The broader reaction from builders like <a href="https://x.com/jakejluo/status/2072380678419705949">@jakejluo</a> and <a href="https://x.com/levie/status/2072519377371459836">@levie</a> was that this pattern will generalize to large-scale document, code, and knowledge workflows.</p></li><li><p><strong>AI-agent evaluation is quickly becoming its own subfield</strong>: <a href="https://x.com/random_walker/status/2072375245969719374">@random_walker</a> noted several new papers advancing agent evaluation and described it as a distinct discipline. Practical examples included <strong>Agent Arena</strong> re-enabling Fable 5 in agent mode <a href="https://x.com/arena/status/2072423538641031372">@arena</a>, <strong>AA-AgentPerf</strong> for agents-per-megawatt system benchmarking <a href="https://x.com/ArtificialAnlys/status/2072254061244825981">@ArtificialAnlys</a>, and <strong>WorldModelGym</strong>, which evaluates whether a world model actually supports good decision-making rather than just producing plausible simulations <a href="https://x.com/RekaAILabs/status/2072325792558956573">@RekaAILabs</a>.</p></li><li><p><strong>There is also a push toward better reporting pipelines for AI failures</strong>: <strong>FLARE-AI</strong>, launched with a coalition spanning cyber and AI safety researchers, aims to standardize <strong>flaw and incident reporting</strong> so issues can be routed to the right developers and registries instead of disappearing into siloed intake forms <a href="https://x.com/ClementDelangue/status/2072401982569025742">@ClementDelangue</a>, <a href="https://x.com/ShayneRedford/status/2072408461015707883">@ShayneRedford</a>.</p></li></ul><p><strong>Systems, Inference, and Architecture Work Worth Watching</strong></p><ul><li><p><strong>NVIDIA&#8217;s TwoTower result stands out as a concrete speed/quality tradeoff on generation architecture</strong>: <a href="https://x.com/NVIDIAAI/status/2072394812301480067">@NVIDIAAI</a> introduced <strong>Nemotron-Labs-TwoTower</strong>, adapting a 30B model into a diffusion-style language model that writes tokens in parallel via a two-copy setup. Claimed result: <strong>2.42&#215; faster generation</strong> while preserving <strong>98.7%</strong> of the original model&#8217;s quality. <a href="https://x.com/LiorOnAI/status/2072402904867365167">@LiorOnAI</a> summarized the trick as reusing a frozen context model plus a trained writer model, avoiding full retraining from scratch.</p></li><li><p><strong>On-device and browser inference continue to benefit from agentic optimization and specialized runtimes</strong>: <a href="https://x.com/googlegemma/status/2072416614188974274">@googlegemma</a> highlighted <strong>WebGPU Gemma 4</strong> running at <strong>255 tok/s on M4</strong>, attributed to kernels written with Fable 5. <a href="https://x.com/andimarafioti/status/2072335408294236164">@andimarafioti</a> demoed a fully open-source realtime voice stack around <strong>Gemma 4 31B</strong> with <strong>Cerebras</strong> inference, aiming as a drop-in alternative to OpenAI&#8217;s realtime API. At the kernel level, Hugging Face&#8217;s kernels library now exposes MiniMax&#8217;s <strong>MSA kernel</strong> <a href="https://x.com/RisingSayak/status/2072277942554841292">@RisingSayak</a>, and Triton-on-Mac drew interest as well <a href="https://x.com/QuixiAI/status/2072345855093289005">@QuixiAI</a>.</p></li><li><p><strong>Architecture research beyond vanilla LLM scaling also surfaced</strong>: <a href="https://x.com/gklambauer/status/2072213633640075366">@gklambauer</a> pointed to <strong>AdaJEPA</strong>, a LeCun-led world-model approach with <strong>test-time adaptation</strong> via latent-state prediction error; <a href="https://x.com/LiorOnAI/status/2072380547603829224">@LiorOnAI</a> summarized <strong>NEO</strong> as learning reusable causal &#8220;programs&#8221; rather than only next-frame prediction; and <a href="https://x.com/ziv_ravid/status/2072402889092616309">@ziv_ravid</a> highlighted &#8220;training in imagination&#8221; as an active paradigm rather than just speculation.</p></li></ul><p><strong>Top tweets (by engagement)</strong></p><ul><li><p><strong>Fable 5 availability dominated technical attention</strong>: <a href="https://x.com/claudeai/status/2072402636813607381">@claudeai: &#8220;Fable 5 is back.&#8221;</a>, <a href="https://x.com/ClaudeDevs/status/2072429181565288665">@ClaudeDevs on rate-limit resets</a>, and <a href="https://x.com/cursor_ai/status/2072403323844428217">@cursor_ai on Fable 5 leading CursorBench</a>.</p></li><li><p><strong>Systems/infra launch with broad reach</strong>: <a href="https://x.com/NVIDIAAI/status/2072394812301480067">@NVIDIAAI on TwoTower&#8217;s 2.42&#215; faster generation at 98.7% quality retention</a>.</p></li><li><p><strong>Open model ecosystem momentum</strong>: <a href="https://x.com/Zai_org/status/2072349453361557898">@Zai_org launching ZCode for GLM-5.2</a> and <a href="https://x.com/vipulved/status/2072321276094673083">@TogetherCompute announcing its $800M Series C at an $8.3B valuation</a>.</p></li><li><p><strong>High-signal tooling and knowledge-layer releases</strong>: <a href="https://x.com/LangChain/status/2072376975545798792">@LangChain/OpenWiki</a> and <a href="https://x.com/cognition/status/2072368168182432109">@cognition/Devin Security Swarm</a>.</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><h3><strong>1. Open-Weight Model Releases and Local Runtime Benchmarks</strong></h3><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-not-much-happened-today-900">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency]]></title><description><![CDATA[The software factory vision met resistance today from speakers defending human understanding and control.]]></description><link>https://www.latent.space/p/aiewf-daily-dispatch-agency</link><guid isPermaLink="false">https://www.latent.space/p/aiewf-daily-dispatch-agency</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Thu, 02 Jul 2026 06:13:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yyfq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yyfq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yyfq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yyfq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg" width="1280" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:666370,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204578515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yyfq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yyfq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5879355c-6a34-432c-bd6a-a4ace5715f5e_1280x850.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">&#8220;You can&#8217;t one-shot design.&#8221; Paul Bakaus at AIEWF today.</figcaption></figure></div><p>Wednesday was autoresearch day on the AI Engineer World&#8217;s Fair main stage.</p><p>Autoresearch is &#8212; you guessed it &#8212; a kind of loop. Introspection co-founder Roland Gavrilescu explained it best in <a href="https://www.latent.space/p/autoresearch-introspection">an interview</a> with Latent Space this morning. He said autoresearch &#8220;allows you to build loops in which agents help maintain the system itself.&#8221; He called it an &#8220;outer loop&#8221; that &#8220;studies and maintains&#8221; the primary, inner loop.</p><p>While autoresearch was not specifically mentioned by Anthropic&#8217;s Thariq Shihipar, who works on Claude Code, his keynote reflected the same idea of continuous discovery and adaptation. &#8220;The models are grown, not developed,&#8221; he said. &#8220;We sort of figure out and learn with the model as we use it.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pnSj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pnSj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pnSj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg" width="1280" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:722340,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204578515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pnSj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pnSj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101677a1-5186-4db9-9b98-8376f22dbad0_1280x800.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Anthropic&#8217;s Thariq Shihipar at AIEWF.</figcaption></figure></div><p>Former Google engineering leader Addy Osmani also spoke about loops, but his framing differed sharply from Gavrilescu&#8217;s.</p><p>Where autoresearch puts agents into the loop that studies and maintains the system, Osmani argued that the outer loop should remain human. &#8220;Agents can run much more of the inner execution loop,&#8221; he said. &#8220;But that outer loop is still engineering.&#8221; His summary was even more direct: &#8220;That inner loop is capability. The outer loop is agency.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hchG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hchG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 424w, https://substackcdn.com/image/fetch/$s_!hchG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 848w, https://substackcdn.com/image/fetch/$s_!hchG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!hchG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hchG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg" width="1280" height="705" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:705,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:306299,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204578515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hchG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 424w, https://substackcdn.com/image/fetch/$s_!hchG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 848w, https://substackcdn.com/image/fetch/$s_!hchG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!hchG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd19ebe6a-9249-4038-a2a1-7ba5ec745724_1280x705.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Addy Osmani&#8217;s Agency Ladder</figcaption></figure></div><h2><strong>Human agency is still important</strong></h2><p>This tension between what agents should do and what human engineers should retain was a recurring theme throughout the day. I also detected some pushback against the &#8220;software factory&#8221; framing that <a href="https://www.latent.space/p/aiewf-daily-dispatch-loops">dominated Tuesday</a>. This tweet from Notion&#8217;s Geoffrey Litt summed it up:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/geoffreylitt/status/2072033306233688511?s=20&quot;,&quot;full_text&quot;:&quot;<span class=\&quot;tweet-fake-link\&quot;>@charlieholtz</span> preach!\n\n&#8220;Factories&#8221; is a depressing vision of the future, metaphors matter &quot;,&quot;username&quot;:&quot;geoffreylitt&quot;,&quot;name&quot;:&quot;Geoffrey Litt&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/722626068293763072/4erM-SPN_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-30T19:03:31.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HMFXT3iaQAAUFln.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/FavLNmrwGs&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:22,&quot;retweet_count&quot;:15,&quot;like_count&quot;:274,&quot;impression_count&quot;:35459,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Litt drew a large audience in the Design Engineering track today, where he <a href="https://x.com/geoffreylitt/status/2072382778763583603">spoke about</a> &#8220;how and why humans need to understand our code.&#8221; Lily Zhang tweeted <a href="https://x.com/lily_gpupoor/status/2072469046000496963">the key takeaway</a>: &#8220;The future will be very polarized: those who understand will keep having the next big idea. Those who delegate understanding will be replaced by the agent.&#8221;</p><p>Later, Litt <a href="https://x.com/geoffreylitt/status/2072525040873328890">posted a thread</a> expanding on his argument. Although he acknowledged that agents are increasingly capable of handling more of the process, humans still need to understand what is happening. &#8220;You can learn what the agent is doing to make sure you can be an active participant in the creative process,&#8221; he wrote.</p><p>Another AIEWF speaker seeking to reinforce human agency was Paul Bakaus, who ran a session about his new design tool, Impeccable. Bakaus rejected both extremes: continuing to design entirely by hand, or &#8220;loop-maxing&#8221; toward a fully hands-off process. &#8220;The truth is somewhere in the middle,&#8221; he told me after his session.</p><p>His goal is to let agents handle the laborious first 80% of the work, before bringing the human back in &#8220;for the last 20% to make it a unique thing &#8212; to really put in your taste, your point of view.&#8221;</p><div class="pullquote"><p>&#8220;There is no auto, and there will be no auto.&#8221;<br>- Paul Bakaus, Impeccable</p></div><p>For Bakaus, that is not simply a temporary limitation of today&#8217;s models. It is also about authorship and accepting responsibility for your work. &#8220;People need purpose, and they want to play a role in whatever they create,&#8221; he said. &#8220;When you work with the agent, then you feel more ownership of the product.&#8221;</p><p>This philosophy is built into Impeccable itself. &#8220;There is no auto, and there will be no auto,&#8221; Bakaus told the audience. What he means is that his product will never &#8220;one-shot&#8221; a solution &#8212; the user must be involved in the design process. &#8220;The point is to give you a way to steer what you want to end up with,&#8221; he added.</p><h2><strong>Generative media</strong></h2><p>The same question surfaced during a panel on generative media. As image, video and audio models become more capable, the issue is not merely what they can generate, but whose judgment shapes the result.</p><p>Nicole Brichtova, who works on Google&#8217;s generative media products, including Nano Banana, drew a distinction between average preference and cultivated expertise. &#8220;Somebody who has honed a craft has a very different level of expertise,&#8221; she said. &#8220;You see things that the average human will not.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JX28!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JX28!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JX28!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JX28!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JX28!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JX28!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg" width="1280" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:746419,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204578515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JX28!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JX28!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JX28!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JX28!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cfc820-15ab-4a95-ac1a-f453700888dd_1280x850.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This matters because every model has a default aesthetic, whether its creators acknowledge it or not. &#8220;It ends up being us,&#8221; Brichtova said. &#8220;It ends up being the modeling teams.&#8221; She suggested that model developers may need to work more closely with people who have &#8220;a really creative point of view&#8221; &#8212; effectively bringing the art director back into the loop.</p><p>Shane Gu made the same point more broadly. Even as models become better at generating and refining their own outputs, he argued, humans must retain the sensitivity to notice what is wrong, generic or insufficient.</p><p>&#8220;Maybe right now the AI can do a lot of all the promptings and it&#8217;s sufficient, but if it&#8217;s like that, never be satisfied [that] AI is generating the content. Always find your sensitivity.&#8221;</p><h2><strong>Agentic sites</strong></h2><p>Even the web itself &#8212; the ultimate human information network &#8212; is grappling with how much automation to use.</p><p>In his session this afternoon on &#8220;agentic sites,&#8221; Adobe principal scientist Carlos Sanchez demonstrated websites that assemble and personalize pages in real time based on a visitor&#8217;s intent. He presented this transition as increasingly inevitable: &#8220;This is now possible. It&#8217;s only going to get better. It&#8217;s only going to get cheaper. It&#8217;s only going to get faster.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OGEi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OGEi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OGEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg" width="1280" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:654399,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204578515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OGEi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OGEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fdeb44-4405-444d-9ff1-977f40b1976f_1280x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>But Sanchez also sounded a note of caution. &#8220;With AI, it&#8217;s very easy to build things, but it&#8217;s hard to know what to build,&#8221; he told me afterwards. That becomes especially important when an agent is generating experiences on behalf of a brand. &#8220;You cannot just generate the whole site,&#8221; he said, because the result may stray outside the brand&#8217;s guidelines.</p><p>That brings the discussion back to autoresearch. Agents may increasingly be able to observe, evaluate and improve other agents, but humans must still define the goals, judge the results, and take responsibility for what the loop produces.</p><p>As impressive as agentic technology is now, and as compelling an idea as automated &#8220;software factories&#8221; might be, you still need humans in the loop.</p>]]></content:encoded></item><item><title><![CDATA[Autoresearch: The feedback loop behind self-improving agents]]></title><description><![CDATA[Introspection co-founder Roland Gavrilescu explains autoresearch, agent &#8220;recipes,&#8221; self-improving loops, and why humans remain central to the software factory.]]></description><link>https://www.latent.space/p/autoresearch-introspection</link><guid isPermaLink="false">https://www.latent.space/p/autoresearch-introspection</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Wed, 01 Jul 2026 23:52:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!p4Th!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!p4Th!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!p4Th!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!p4Th!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:655460,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204548385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!p4Th!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!p4Th!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe316b2cb-4200-4a71-bdcc-c398467b53ef_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Introspection&#8217;s Roland Gavrilescu at AIEWF.</figcaption></figure></div><p>We&#8217;ve heard a lot about loops at the AI Engineer World&#8217;s Fair this week. Another buzzword is <strong>autoresearch</strong>, which involves building an &#8220;outer loop&#8221; where agents help maintain and improve the primary system, using feedback signals, evals and human input to make progress over time.</p><p>At least, that was the framing of <a href="https://x.com/rolandgvc">Roland Gavrilescu</a>, co-founder and CEO of Introspection &#8212; a new company building infrastructure for deploying these self-improving systems. Before starting the company, Gavrilescu worked on agent infrastructure and cloud agents at xAI, where he met his co-founder, <a href="https://www.linkedin.com/in/julianbright/">Julian Bright</a>.</p><p>Ahead of his &#8220;Autoresearch in the Wild&#8221; session at the AI Engineer World&#8217;s Fair today, I spoke with Gavrilescu about the shift from agent harnesses to feedback loops, the role of the open-source Pi framework, and why autonomous software factories must first learn from humans.</p><h2><strong>From xAI to Introspection</strong></h2><p><strong>Latent Space:</strong> How did your new company, Introspection, come about?</p><p><strong>Roland Gavrilescu:</strong> Last year, I was at xAI, where I met my co-founder. We were working on agent infrastructure and cloud agents, and we felt there was a new agent form factor that needed to be explored further. xAI was not necessarily the environment where we could focus completely on that.</p><p>We decided to leave and ask what a company designed around this new form factor might look like. We were interested in what made companies such as Cursor and Cognition successful, and how we could turn some of those ideas into a product that others could use.</p><p>That became the basis for Introspection.</p><p>Autoresearch allows you to build loops in which agents help maintain the system itself. The challenge is designing the right signals and feedback mechanisms so agents can improve the system, make architectural decisions and move in the right direction without constantly being bottlenecked by humans.</p><h2><strong>The loop becomes the product</strong></h2><p><strong>Latent Space:</strong> Your session is titled &#8220;Autoresearch in the Wild&#8221; &#8212; what will it cover?</p><p><strong>Gavrilescu:</strong> We have heard a lot about what autoresearch can do for improving experiments, but we wanted to talk about what these loops look like in production.</p><p>We are presenting three patterns that we think form the basis of a new blueprint.</p><p>The first is that <strong>the loop is the product</strong>. We have moved from focusing on models, to harnesses, and now to loops. The key question is whether you can define the right feedback mechanisms so agents can take on more work without generating more slop.</p><p>The second pattern concerns what the loop generates and how you track it over time. We are proposing a concept called an agent <strong>recipe</strong>.</p><p>We moved from agent tools to agent skills. Recipes are a larger container that brings together the components needed to encode human expertise: evals, judges, signal processing and the information that feeds back into the loop.</p><p>The goal is to create a portable format that agents can iterate on, almost like a research laboratory, but in a provider-agnostic way.</p><p>The third pattern is about what we optimize for. How can the system become both better and cheaper over time?</p><p>Companies such as Cursor and Cognition have shown that these products can work. The next stage is making them more accessible, faster and cheaper, and gradually distilling the capabilities of frontier models into systems that you own and that are customized for your environment.</p><h2><strong>Agent recipes</strong></h2><p><strong>Latent Space:</strong> Can you explain more about what an agent recipe is&#8230;</p><p><strong>Gavrilescu:</strong> It&#8217;s like a description of the ingredients you need and how they evolve.</p><p>The idea comes partly from data recipes used in model post-training. A data recipe describes how much data from different domains should be baked into a model.</p><p>Agent recipes are similar. A recipe might describe how your harness works with different models, the evals you use, the judges you have created, the human expertise you have captured and the failures that led to new evals.</p><p>Imagine that tomorrow you suddenly gained access to the Devin codebase. The code alone would not necessarily be that helpful if you could not see how the team arrived at the current version. You would want to understand the failures, mistakes and decisions that informed it.</p><p>A recipe captures that process. You begin with a baseline and then record how each signal produced a new judge, embedded new human expertise or led you to introduce a different model.</p><h2><strong>The inner loop and the outer loop</strong></h2><p><strong>Latent Space:</strong> Does autoresearch mean orchestrating multiple agents, or can it involve one agent repeatedly working and verifying its results?</p><p><strong>Gavrilescu:</strong> You can think of the system as having an inner loop and an outer loop.</p><p>The inner loop is the primary system interacting with users and performing the work. Autoresearch is more concerned with the outer loop: another system that studies and maintains the primary system.</p><p>The question is how to design that outer loop so it makes progress on the right problems without consuming an unreasonable number of tokens while deciding what to do.</p><h2><strong>Pi as the Linux of agent harnesses</strong></h2><p><strong>Latent Space:</strong> You have compared Pi to Linux. In that analogy, is Introspection something like Red Hat?</p><p><strong>Gavrilescu:</strong> Pi is like the Linux of agent harnesses. Linux has distributions such as Ubuntu, but the underlying system is designed to be extended. Pi is similar: it was never intended to be run as an unchanged, vanilla product. Pi separates the agent loop from its extensions and configuration, which makes the agent portable. You can spin up several different agents by loading different files into the runtime.</p><p>We saw an opportunity to combine that extensibility with recipes and open-source building blocks that can evolve for each customer while remaining portable and easy to deploy.</p><h2><strong>Making loops reliable in production</strong></h2><p><strong>Latent Space:</strong> Reliability and the messy reality of agent loops have been recurring themes at the conference. How does Introspection address those problems?</p><p><strong>Gavrilescu:</strong> The product is designed around the point at which you are ready to move into production.</p><p>You need to know what infrastructure is required to make the loops work, keep costs under control and maintain security. The managed infrastructure covers what is necessary for these systems to operate in production.</p><p>A major part of our focus is bringing the kind of infrastructure available inside frontier AI laboratories to a product that other companies can deploy.</p><h2><strong>Humans remain part of the system</strong></h2><p><strong>Latent Space:</strong> What about the human in the loop?</p><p><strong>Gavrilescu:</strong> These loops are designed with humans in the loop because you need the right signals as the system makes progress.</p><p>The human can effectively become a tool and a source of signals. Agents can be trained to ask people questions through an &#8220;ask a human&#8221; tool.</p><p>During its first few loops, an agent may rely heavily on asking questions and learning what a human would do. Over time, it accumulates those preferences and can become increasingly autonomous.</p><p>It is similar to an employee joining a new company. Initially, that employee asks a lot of questions. As they learn how the organization works, they can make more decisions independently.</p><h2><strong>Taking agent infrastructure into vertical markets</strong></h2><p><strong>Latent Space:</strong> So what kinds of use cases are you seeing?</p><p><strong>Gavrilescu:</strong> We are concentrating on vertical agents.</p><p>Coding agents are clearly working, and we have seen a number of companies succeed in that area. The next question is how to deploy agents in vertical and non-coding domains.</p><p>Companies in those markets are asking how they can do this securely without becoming dependent on a single provider. They want the deployment to belong to them, they want to retain ownership of their data, and they do not want to be locked into OpenAI or Anthropic. Introspection is intended to provide infrastructure that addresses those requirements using open-source building blocks.</p><p>Frontier AI labs have developed sophisticated internal agent technology. We want to bring similar capabilities into vertical SaaS and services businesses.</p><h2><strong>Why the work happens in Git</strong></h2><p><strong>Latent Space:</strong> Is Introspection mainly intended for developers, or will product managers and other business users work with it?</p><p><strong>Gavrilescu:</strong> We are initially focusing on software engineers in vertical SaaS companies.</p><p>We want the environment to be agent-friendly, meaning agents can work inside their own repositories and codebases. Everything is Git-based, and Git becomes the audit log that you maintain over time.</p><p>In the future, there will be interfaces that enable product managers and others to participate. But we are already seeing product managers move closer to code.</p><p>We think the right initial form factor is a human-to-agent interface in which the actual work and its history live in Git.</p><h2><strong>From orchestras to software factories</strong></h2><p><strong>Latent Space:</strong> Does Introspection fit within the broader idea of software factories?</p><p><strong>Gavrilescu:</strong> Yes. Designing the loops is essentially designing the factory. The remaining question is how much autonomy the factory should have.</p><p>There has also been discussion about &#8220;orchestras, not factories.&#8221; That distinction is really about the level of autonomy.</p><p>An orchestra might retain a human conductor who controls how the loops operate. A factory implies something more fully autonomous.</p><p>But you should build toward the factory rather than assume you can create a completely autonomous factory on the first day. Models do not initially possess all the context or understand every decision people inside an organization make. You cannot simply capture all of that knowledge in a Markdown file.</p><p>The right approach is to design the human as a core component of the factory. The early system should extract tacit knowledge and workflows from people over time, rather than attempting to automate everything immediately.</p><h2><strong>How to start with autoresearch</strong></h2><p><strong>Latent Space:</strong> What would you recommend to engineers who want to experiment with autoresearch?</p><p><strong>Gavrilescu:</strong> The first step is to invest in your signals. What are the things you actually want agents to respond to?</p><p>Product feedback is a good example. Not all feedback carries the same value, and you cannot respond to every individual data point. You need a mechanism for filtering the signals and identifying which ones an agent should act on.</p><p>The second requirement is control over cost. You do not want to wake up to an unexpected thousand-dollar bill because an agent has been running an inefficient loop.</p><p>The third is to follow the research. Look at the kinds of harnesses models are being trained to use and remain close to those patterns. Study how research labs use data recipes and consider how those ideas can be applied to your own product.</p><p>The broader goal is to turn your product organization into a miniature research lab, with agents acting as miniature researchers.</p>]]></content:encoded></item><item><title><![CDATA[How Cursor deploys AI inside the enterprise]]></title><description><![CDATA[Cursor's Pauline Brunet explains how her team of Forward Deployed Engineers help organizations implement agents &#8212; essentially setting up software factories.]]></description><link>https://www.latent.space/p/cursor-forward-deployed-engineers</link><guid isPermaLink="false">https://www.latent.space/p/cursor-forward-deployed-engineers</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Wed, 01 Jul 2026 19:03:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!e2BU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!e2BU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!e2BU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!e2BU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:678489,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204513174?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!e2BU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!e2BU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8c541c-264c-476f-b47c-029cd970acf9_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Pauline Brunet, VP of Forward Deployed Engineering at Cursor, at AIEWF.</figcaption></figure></div><p>Forward deployed engineering has quickly become one of the most prominent roles in enterprise AI. Sitting somewhere between software engineering, product development and customer implementation, forward deployed engineers [FDEs] work directly with organizations to implement AI capabilities.</p><p>At Cursor, the role is especially ambitious. <a href="https://www.linkedin.com/in/pauline-brunet/">Pauline Brunet</a>, the company&#8217;s VP of Forward Deployed Engineering, is building a team that works with organizations to implement agents across the entire software development lifecycle.</p><p>In an interview with Latent Space at the AI Engineer World&#8217;s Fair, Brunet discussed Cursor&#8217;s vision of an &#8220;AI software factory,&#8221; the challenge of expanding agent adoption beyond individual enthusiasts, and what engineers need to demonstrate if they want to move into forward-deployed work.</p><h2>What forward deployed engineering means at Cursor</h2><p><strong>Latent Space:</strong> To begin with, how do you define forward deployed engineering?</p><p><strong>Pauline Brunet:</strong> Forward deployed engineering depends on the business, the product, and the customer. You have to consider how configurable the application is. Is it something customers can use out of the box, or are you deploying something complex and highly configurable?</p><p>You also have to consider where customers are in their journey.</p><p>I don&#8217;t think of forward deployed engineering as a team that supports a traditional, out-of-the-box deployment. I think of it as a team that goes on-site, works inside a customer&#8217;s systems and tools, and deploys applications or platforms that help solve challenges at scale.</p><p>Those deployments are highly configurable and customized around the customer&#8217;s workflows, processes, systems, and tools.</p><p><strong>Latent Space:</strong> Cursor&#8217;s customers are predominantly engineers. How does the FDE role apply to the way they use the product?</p><p><strong>Brunet:</strong> Cursor is an AI coding platform and coding assistant. We work with people on AI-assisted coding, synchronous and asynchronous agents, and ultimately the idea of an AI software factory.</p><p>Today, we work with customers across many industries, including financial services, telecommunications, software development, technology, and semiconductors.</p><p>We help transformation leaders, IT leaders, and CTO organizations create an AI software factory across their operations. That includes how they plan and design software, how they write code, how they test and review it, and how they deploy and maintain applications at scale. So, very focused on the software development lifecycle from start to finish.</p><h2>Building Cursor&#8217;s FDE team</h2><p><strong>Latent Space:</strong> How large is Cursor&#8217;s FDE team?</p><p><strong>Brunet:</strong> We&#8217;re growing rapidly. Our goal is to grow the team tenfold by the end of December.</p><p><strong>Latent Space:</strong> Are your current FDE employees primarily engineers, or does the team also include product specialists?</p><p><strong>Brunet:</strong> They are all engineers. We hire software engineers with at least five years of experience and extensive customer-facing experience.</p><p>These are people who have developed and shipped code in production. They have built and designed systems, and they can make trade-off decisions and evaluate which systems or technologies should be used.</p><p>They also need customer-facing experience. We have people who previously worked at companies including Spotify, Rippling, and Palantir, and who have deployed production systems for customers.</p><h2>From coding assistants to software factories</h2><p><strong>Latent Space:</strong> You mentioned the term &#8220;<a href="https://www.latent.space/p/software-factories">software factory</a>,&#8221; which has begun appearing more frequently in the industry. What does that term mean to Cursor?</p><p><strong>Brunet:</strong> For Cursor, it is about the software development lifecycle from start to finish: how you plan, design, write, review, test, and deploy code.</p><p>Today, those stages are often handled by different teams. You might have a design team, a development team, and a product manager working alongside them. Each group may be optimizing its own work with AI-assisted coding, but the process remains siloed.</p><p>We want to help customers across the entire lifecycle. You should be able to say, &#8220;Here is the feature I want to develop,&#8221; and then have long-running agents work with you across every step. That could include creating the plan and product requirements document, producing a demonstration of what the feature might look like, writing and testing the code, putting it into production, and maintaining it.</p><p>Issues and product feedback should also feed back into that same lifecycle. For us, a software factory means long-running agents helping people throughout that entire process.</p><p><strong>Latent Space:</strong> So it is broader than agent orchestration alone?</p><p><strong>Brunet:</strong> Correct. Exactly.</p><h2>Moving beyond individual AI adopters</h2><p><strong>Latent Space:</strong> What problems are enterprises encountering as they try to implement agent technology?</p><p><strong>Brunet:</strong> One challenge is that adoption is still concentrated among early adopters.</p><p>Within an organization, you might have 10% or 20% of people who are enthusiastic early adopters. They have done great work using local agents and cloud agents for their own tasks, and they have become highly productive.</p><p>What is missing in the next phase is the ability to use long-running agents across teams, processes, and workflows.</p><p>That requires more support from the top of the organization. Leadership has to say, &#8220;This is a priority, and this is how we want to automate or change this process.&#8221;</p><p>For the FDE team, it is therefore important to find the right champions inside an organization: people who want to meaningfully change the business and who will work with us and their internal teams to transform how work gets done.</p><h2>Standardizing work with cloud agents</h2><p><strong>Latent Space:</strong> <a href="https://www.latent.space/p/ahmad-osman-local-ai">Local AI</a> appears to be gaining momentum, partly because of the increasing availability of open-source models. Are you doing more local AI implementation work with customers?</p><p><strong>Brunet:</strong> We have local agents that people run through the desktop application or the CLI, and that experience is largely self-service. People have adopted the technology at a phenomenal rate, particularly across Cursor&#8217;s user base.</p><p>We are also seeing people adopt cloud agents because they are excited about being able to run tasks without keeping their laptops half open. Agents can now work in the cloud on tasks that previously ran locally.</p><p>What becomes interesting is when this moves beyond an agent helping with one person&#8217;s job. The next question is how agents can work across a function, team, or organization so that processes are automated consistently. For example, you could have a QA agent applying the same process across several development teams.</p><p>We are receiving a lot of questions from customers about those kinds of use cases.</p><h2>How customer deployments influence Cursor&#8217;s roadmap</h2><p><strong>Latent Space:</strong> Do the lessons from these deployments feed back into the core Cursor product?</p><p><strong>Brunet:</strong> Yes. The forward deployed engineering team works very closely with customers on their use cases, so we are naturally a good way for the product and engineering teams to understand what customers want to build next.</p><p>We work closely with those teams and play a significant role in helping shape Cursor&#8217;s product roadmap.</p><h2>The changing role of the forward deployed engineer</h2><p><strong>Latent Space:</strong> As agents become more autonomous, how do you expect the FDE role to evolve?</p><p><strong>Brunet:</strong> I think the role is going to change drastically. I always say that if we are doing the same job we were doing six months ago, we have done something wrong.</p><p>Right now, people are still looking for inspiration about the use cases they can solve, so we want to propose new possibilities.</p><p>In software development, for example, we can show how designers and product managers might work seamlessly in Cursor alongside developers and testing teams.</p><p>We might also ask whether a company has considered using long-running agents to handle call-center or ticketing processes from start to finish.</p><p>As we work across industries such as healthcare, life sciences, the public sector, retail, and consumer packaged goods, we will continue identifying use cases across marketing, sales, and supply-chain operations. The FDE role will evolve alongside those possibilities.</p><h2>How engineers can prepare for an FDE career</h2><p><strong>Latent Space:</strong> There are around 7,000 AI engineers at this conference. What advice would you give developers who want to move into forward deployed engineering?</p><p><strong>Brunet:</strong> I&#8217;ve had this conversation five or six times already today. We are looking for builders with software engineering experience: people who have identified a problem and built a production-grade application or system from start to finish.</p><p>You should have designed it, developed it, tested it, and put it into production with real users.</p><p>My recommendation is to find those kinds of projects inside your organization and take ownership of them from beginning to end. Make sure you understand why you made each design decision.</p><p>How did you select the database? How did you choose the different services? Why did you design the system in that particular way? What were the trade-offs?</p><p>You should also understand the measurable return on investment, both in traditional business terms and through evaluations that demonstrate the value you are creating for internal customers.</p><p>If you want to get into forward deployed engineering, become familiar with these kinds of projects, gain experience delivering them, and learn how to explain the decisions you made.</p>]]></content:encoded></item><item><title><![CDATA[🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI]]></title><description><![CDATA[Why the Llama lead left Meta for drug discovery, PEARL's zero-shot OpenBind win, and what becomes possible when co-folding finally crosses the accuracy threshold.]]></description><link>https://www.latent.space/p/the-coolest-diffusion-research-isnt</link><guid isPermaLink="false">https://www.latent.space/p/the-coolest-diffusion-research-isnt</guid><dc:creator><![CDATA[Brandon Anderson]]></dc:creator><pubDate>Wed, 01 Jul 2026 14:42:39 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/204393300/c3e6a250afb2434055b308d47570a862.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>This episode has a fun personal twist: There&#8217;s a counterfactual world where I was employee #1 at  <a href="https://www.genesis.ml/">Genesis Molecular AI</a>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> the company behind today&#8217;s episode. A certain introduction happened a few weeks too late and I had already happily signed at Atomwise<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, another ML-for-drug-discovery startup. Same problem, different company. I was certain ML was going to transform small molecule drug discovery. Early results were underwhelming. Useful at times, but nowhere near revolutionary. In the last year I&#8217;ve seen signs that ML is finally ready to deliver on my convictions from a decade ago. Genesis is one of the places that might have finally cracked this problem. I was super excited to come full circle and catch up with co-founder <a href="https://www.linkedin.com/in/evanfeinberg">Evan Feinberg</a> and CTO <a href="https://www.linkedin.com/in/edunov/">Sergey Edunov</a>.</p><p>If you are at all interested in small molecule drug discovery, we think you will find this fascinating!</p><p>In our nearly two hour chat we cover:</p><ul><li><p>What is small molecule drug discovery, and why is it hard</p></li><li><p>Structure prediction as a hotbed of innovation in AI algorithms</p></li><li><p>How advances in AI elsewhere have enabled stepwise improvements in predictive power</p></li><li><p>How the community benchmarks are essentially calling AI slop good enough</p></li><li><p>The Genesis flagship model (PEARL) can routinely hit a threshold that is necessary for real-world applications</p></li><li><p>New agentic workflows enabled by these highly accurate models</p></li></ul><p>Read on for more, and also some personal thoughts on the future at the end.</p><h1>The coolest diffusion research is happening at Genesis</h1><p>Sergey Edunov came to Genesis from Meta where he led Llama 2 training and Llama 3 pretraining. Sergey was a former physicist who thought he was done with physics after many years of training LLMs. Then, he discovered Genesis, and was blown away with all the novel architecture work they&#8217;ve been developing.</p><p>It probably surprises no one that modern LLM research has not resulted in fundamentally novel or exciting updates in architectures since almost the advent of the transformer &#8212; the entire field is using variants on the same idea that came out in the original &#8220;Attention is all you need&#8221; paper. Sure, some were quite useful (mixture-of-experts in particular allowed for the massive model paradigm we&#8217;re at today), but there was very little conceptually exciting.</p><blockquote><p>&#8220;We sort of had to wait for the right primitive to get created, and that turned out to be diffusion&#8230; Actually, some of the most innovative diffusion research that&#8217;s happening in our field is happening in 3D structure prediction right now.&#8221; &#8212; Evan Feinberg</p></blockquote><p>The field of 3D structure prediction on the other hand has been a hotbed of research. Genesis&#8217; recent model <a href="https://www.genesis.ml/news/introducing-pearl">PEARL</a> (Place Every Atom at the Right Location) is able to understand protein flexibility, and model not just where the ligand goes, but also make small adjustments of the protein so that the two fit better than either alone. The field knew this was missing for a long time, but it was really hard to model until now.</p><h1>Agentic Discovery</h1><p>What makes this problem so hard? As Sergey points out, there are 10^60 possible drug-like small molecules. You&#8217;ll never be able to search them all, and trying to find the good ones is something like finding a needle in a haystack &#8212; except everything except your needle is dangerous.</p><blockquote><p>&#8220;There are 10 to the 60 drug-like small molecules in the universe&#8230; it&#8217;s like finding a needle in a haystack, where everything except your needle is very, very dangerous.&#8221; &#8212; Sergey Edunov</p><p>&#8220;Or finding hay in a needle stack might be a more apt analogy.&#8221; &#8212; Evan Feinberg</p></blockquote><p>Trying to solve the multi-parameter optimization problem is even worse. What makes a strong binder and a molecule with good &#8220;ADMET Properties&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> are oftentimes at tension with each other. For example, a good binder is likely greasy, but a greasy molecule is likely insoluble so it won&#8217;t enter the bloodstream and get to where it needs to go!</p><p>Genesis&#8217; advances in generative AI have now pushed them beyond the threshold where they believe agentic drug discovery loops are finally possible. We all remember the early days of LLMs. They were great chatbots but terrible agents, as small errors compounded rapidly into uselessness. As LLMs got better, the usefulness of agents rapidly improved. Evan and Sergey argue that their models at Genesis recently passed a similar threshold. Their internal agentic drug-discovery system (code named SAPPHIRE) can now iterate like a chemist: look at and reason about poses, form hypotheses, read literature, use internal tools, create candidates for the next iteration. Combining this with automated lab partnerships like the one Genesis has with <a href="https://incyte.com/">Incyte</a>, we&#8217;re rapidly approaching a time of drug discovery agents running 24/7 making/testing new molecules. Exciting times!</p><h1>Benchmark crisis: Everyone&#8217;s favorite benchmark is slop</h1><p>One surprising point that isn&#8217;t talked enough about: the academic field of &#8220;co-folding&#8221; has settled on a benchmark value of &#8220;2 Angstrom RMSD&#8221; as a metric for a &#8220;good pose&#8221;. Evan does not mince words: this threshold is just bad. Perhaps even deceptively bad. For many strong binders, there&#8217;s a very clear pose, one that you can even directly resolve in the PDB electron density! And yet, with a 2&#197; RMSD threshold, you can get the pose quite wrong in ways that might even mislead a medicinal chemist. For example, flip around an aromatic ring, and everything looks reasonable, but you&#8217;re no longer modeling the right interactions.</p><p>Evan makes the strong claim that 1&#197; RMSD is really the threshold necessary to ensure the core of the molecule is sitting where it needs to be, and models all interactions.</p><blockquote><p>&#8220;If your model is sitting at 1.8, 1.9 Angstrom RMSD, that&#8217;s slop, most likely.&#8221; &#8212; Evan Feinberg</p></blockquote><p>As a simple example, he points out hydrogen bonds which are responsible for many of the most important interactions in protein-ligand systems. Hydrogen bonds only have a 0.6&#197; range to be valid! Clearly if you&#8217;re accurately resolving all H-bonds, you generally have to be doing much better than the 2&#197; threshold.</p><p>This is clearly a hard-fought lesson for Evan and Genesis. In their opinion, the community is stuck on these benchmarks because academics developing methods were not users. Evan does see signs of life, with the use of new metrics such as lDDT for co-folding. Hopefully soon the community can agree that &#8220;1.8&#197; RMSD is slop&#8221;, and start hill climbing on this much harder task.</p><p>For a more thorough exploration of the weaknesses in conventional benchmarks, see the <a href="https://arxiv.org/abs/2510.24670">PEARL technical report</a>.</p><h1>PEARL tops OpenBind</h1><p>Which makes what happened next all the more striking. Near the end of the podcast, we talked about a recent &#8220;proof-is-in-the-pudding&#8221; moment for Genesis &#8212; evaluating their <a href="https://www.genesis.ml/news/zero-shot-pearl-system-surpasses-all-cofolding-models-on-openbind">PEARL model</a> on a recently released OpenBind benchmark. This benchmark featured 802 never before seen co-complexes on a target protein EV-A71. This target seems almost custom-chosen to give most classical docking methods a problem. When a ligand binds to the main binding site, the protein moves around to close off the path the ligand used to enter the binding pocket. This process, known as &#8220;induced fit&#8221; is notoriously hard for traditional methods to model. The tradeoff is easy to understand: treating the protein as a static structure, it becomes difficult to place a ligand in a binding pocket. Treat the protein as dynamic, and now you have to simulate complicated processes that take a long time to resolve.</p><p>PEARL was able to model the induced fit of the ligand without running long MD simulations. Across the different evaluation metrics, PEARL came out not just ahead, but oftentimes well ahead of any public model. A truly impressive result.</p><blockquote><p>&#8220;Where PEARL was exceptionally good is figuring out how to move this loop. We are basically correct for every single pose.&#8221; &#8212; Sergey Edunov</p></blockquote><p>Even more exciting, this was done without any fine-tuning, or using any data on the target or homologous targets &#8212; the template PDB was released after PEARL&#8217;s training cutoff.</p><h1>Where does co-folding go now?</h1><p>As someone who has followed or participated in ML techniques for protein-ligand interactions for almost a decade, I was genuinely impressed with the results that Genesis has released recently. This has been many years in development, and I&#8217;m sure Evan and the team had many sleepless nights trying to get to this point. I also think other teams are making similar progress &#8212; both Isomorphic and Deep Origin have released results that seem spiritually similar and combine computation, wetlab data, ML, to achieve genuine predictive power that seemed impossible a decade ago. Sadly, all of the above are closed source so there&#8217;s no way to honestly compare them. Looking at the results I think there might be a time in the not so distant future where we can consider protein-ligand binding &#8220;solved&#8221;.</p><p>I sincerely hope that the academic community can take inspiration from these developments. Once you know something can be done, it&#8217;s much easier to execute. Still, I believe that the key enabler in all of the above was the tight integration of ML, large-scale computation, and real-world drug discovery applications. Sadly academia is just not structured in a way that makes such a development easy.</p><p>With those parting thoughts, we hope you give the podcast a listen!</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>At the time called Genesis Therapeutics</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Now called Numerion</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p> ADMET stands for Absorption, Distribution, Metabolism, Excretion, and Toxicity. This set of about 30 properties all need to be optimized in order for a molecule to be considered a &#8220;good drug&#8221;.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Warp CEO Zach Lloyd on why software factories are the next phase of coding]]></title><description><![CDATA[Warp's founder thinks every major software project will soon run on an automated factory. He discusses why and how engineers should prepare for this shift.]]></description><link>https://www.latent.space/p/software-factories</link><guid isPermaLink="false">https://www.latent.space/p/software-factories</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Wed, 01 Jul 2026 14:28:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kQB7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kQB7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kQB7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kQB7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:777957,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204445868?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kQB7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kQB7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4140e59-9bc8-4685-8af0-0cf86b6f998f_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Warp founder Zach Lloyd in the AI Engineer World&#8217;s Fair expo hall.</figcaption></figure></div><p>I&#8217;ve been covering Warp for a couple of years now, and its rapid evolution from a command-line interface tool to a <a href="https://www.latent.space/p/aiewf-daily-dispatch-loops">software factory</a> platform has been fascinating to watch. The company began in the pre-ChatGPT days, in mid-2021, as a Rust-based terminal. Then when AI hit, it turned into <a href="https://ricmac.org/2025/02/26/warp-launches-ai-first-native-terminal-app-for-windows/">a terminal with integrated coding agents</a>.</p><p>But the competition among CLI tools has dramatically increased in recent years, including from Claude Code, Codex CLI, and Gemini CLI &#8212; three products backed by massive tech companies. This likely led to Warp&#8217;s decision to <a href="https://www.warp.dev/newsroom/2026/4/28/warp-open-sources-its-agentic-development-environment">open-source its core CLI tool</a> in April this year.</p><p>I&#8217;m a Warp user myself, finding it a much more sophisticated tool than my native Mac CLI. But I also admire the company&#8217;s ability to adapt to the times &#8212; a trait I spotted in CEO Zach Lloyd during my first interview with him a couple of years ago. So I was keen to catch up with him at the AI Engineer World&#8217;s Fair this week, where he presented a keynote session on software factories, the new term for orchestrating a team (ahem, a <em>factory</em>) of agents.</p><p><a href="https://www.warp.dev/">Warp</a> has a new agent orchestration platform called <a href="https://www.warp.dev/oz">Oz</a>. It&#8217;s the company&#8217;s answer to what Lloyd believes is an industry transition, from engineers working interactively with agents to automated systems that continuously triage, implement, review, verify and monitor software changes. Oz is intended to connect multiple models and coding harnesses across local environments and isolated cloud sandboxes, while fitting into tools developers already use.</p><p>I spoke to Lloyd just after he made his presentation on-stage, which you can <a href="https://www.youtube.com/live/htM02KMNZnk?si=uhrS4ZX2COBbTem9&amp;t=18296">view on YouTube</a> &#8212; it&#8217;s a good primer to what software factories are. In our one-on-one discussion, we get into the reasons Warp made its software factory pivot, how Lloyd came up with the term (independently, it seems, from similar companies &#8212; like <a href="https://factory.ai/">Factory</a>), and why he expects most significant software projects to operate some form of automated factory within the next year.</p><h2>From individual agents to an automated development loop</h2><p><strong>Latent Space:</strong> When did you first come across the term &#8220;software factory,&#8221; and what attracted you to the concept?</p><p><strong>Zach Lloyd:</strong> I can&#8217;t remember exactly when I started conceiving of it in those terms, but it was within the last six months, as the ability to automate software development became more complete.</p><p>We started with more one-off automation: run an agent in the cloud. A lot of platforms began there. Then it became: run an agent in the cloud on a timer.</p><p>The next question was, what is the most valuable loop to automate? The answer is basically the main loop of software engineering: triage, specification, implementation, review, verification, shipping and monitoring.</p><p>We began building toward this cloud-automation vision about a year ago, before we started building Oz. Over the past few months, the industry has also begun coalescing around the &#8216;factory&#8217; term. There is an entire software-factory track at this conference.</p><p>It is literally what we are gearing our product around. In the next version of Oz, you will set up your factory, see what it looks like and manage the factory floor.</p><p>But I don&#8217;t care that much whether the term sticks. The essential shift is from interactive development to automated development. &#8220;Factory&#8221; is a useful metaphor for that.</p><h2>Building the factory around existing workflows</h2><p><strong>Latent Space:</strong> In your presentation, you showed a software-factory stack containing several of your own products. Is Warp&#8217;s plan to provide the tools that make up that stack?</p><p><strong>Lloyd:</strong> Yes. When you enter Oz, our cloud-agent platform, you will be walked through setting up a factory.</p><p>You choose your repositories, the parts of the software lifecycle you want to automate, and the points where humans should be brought into the loop. Different organizations and codebases will have different preferences. Do you fully automate code review? Do you have humans review certain high-risk changes?</p><p>The system then starts creating the loop. It might pull issues from Jira or Linear, let people submit them through Slack or Teams, and allow developers to redirect an agent from GitHub.</p><p>What is interesting from a product perspective is that most of the factory is not necessarily a new interface. It is an integration into people&#8217;s existing workflows. That is how we are conceiving it, at least.</p><h2>Why Warp is moving beyond the terminal</h2><p><strong>Latent Space:</strong> When I first wrote about Warp, it was building a modern terminal. Code is still important now, but increasingly it is being produced by agents. It looks like Warp has broadened its product vision accordingly...</p><p><strong>Lloyd:</strong> One hundred percent. A good way to think about it is that the company&#8217;s mission has stayed the same since we founded it. It has always been about empowering developers and companies to ship better software more quickly.</p><p>The product has evolved tremendously. It began as a modern version of the terminal, before the current AI wave. The next iteration was a terminal with agents built into it, which we are still investing in and which we have now open-sourced.</p><p>But the world keeps changing. The underlying AI improves so quickly that my view of the future is what I described in the talk: the interactive component is going to become less important.</p><p>As a company, you will want a central place where software gets built and where you can measure the efficiency of that process. I&#8217;m not afraid to redirect what the product becomes. As the underlying technology gets better, companies that do not adapt are going to be left behind.</p><h2>Factory engineering as a new discipline</h2><p><strong>Latent Space:</strong> The word &#8220;factory&#8221; may be off-putting to some developers, given its connotations with mechanism and rote work. What feedback have you received from AI engineers about this pivot?</p><p><strong>Lloyd:</strong> The concept resonates strongly with the economic buyer &#8212; the person running the engineering team.</p><p>For an individual engineer, it can sound mechanized and uncreative. They may think: &#8220;I enjoy coding. Why would I want to work in a factory?&#8221;</p><p>One point I tried to communicate in the talk is that this will become a new engineering discipline. I think it can be extremely interesting if you view the job as meta-engineering: building the system that builds the product.</p><p>It uses many of the same problem-solving skills. You are asking why an agent performs one task well and another poorly. How should you adjust its feedback? What context does it need? How should the workflow change?</p><p>But, for better or worse, the power of these systems and their ability to accelerate software development are so great that writing everything by hand is not going to make sense for much longer.</p><h2>Where forward-deployed engineers fit</h2><p><strong>Latent Space:</strong> Another trend at the conference is <a href="https://www.latent.space/p/forward-deployed-engineers-aiewf">forward-deployed engineering</a>, which often combines aspects of product management, consulting and traditional engineering. How does that fit into the software-factory model?</p><p><strong>Lloyd:</strong> Standing up a software factory potentially involves integrating with a large number of existing systems, depending on the company.</p><p>The factory will work most effectively when it has context from those systems and is integrated throughout the organization&#8217;s workflow. A lot of forward-deployed engineering work in this area is effectively a transformation project.</p><p>It requires real engineering from someone who understands how to configure and deploy one of these systems. We do some of that, and some of our competitors do as well.</p><p>I don&#8217;t know what the final state will look like. Warp is approaching it more as a platform business than a services business. But there is certainly a business today in sending smart people into a company to transform its workflow using these products.</p><h2>Warp as the test bed for Oz</h2><p><strong>Latent Space:</strong> I use Warp as my terminal, including for some coding tasks. What happens to the original Warp CLI product in the software-factory era?</p><p><strong>Lloyd:</strong> When we open-sourced Warp, we put the repository under the control of Oz. We built a software factory around the open-source project, using our own factory platform.</p><p>We are still trying to improve Warp as much as possible. We are doing it with the community, and we are doing a lot of it with agents. In that sense, Warp is a test bed for the factory concept.</p><p>But it is also a product used by almost a million developers, many of whom rely on it as their primary development environment. We use it constantly ourselves, and we still have internal engineers whose job is to improve it. We are simply approaching that work with a factory mindset.</p><h2>Gradual automation, not an overnight replacement</h2><p><strong>Latent Space:</strong> What do you expect the next year to look like, in terms of adoption of software factories?</p><p><strong>Lloyd:</strong> This will not happen all at once. Engineers are not going to wake up one morning and discover that a software factory has replaced their jobs.</p><p>Companies will start with specific use cases, certain types of issues or lower-risk repositories. Those are places where they may be comfortable not having a human review every single line of code.</p><p>They will see how it performs. Then the engineering challenge becomes: instead of merging 20% of pull requests automatically, can we get to 30%, 40%, 50% or 60%?</p><p>There will still be a remaining percentage of work done by people because it is too difficult, ambiguous or dependent on greenfield thinking.</p><p>But I think this shift will happen over the next year. My prediction is that every significant software project will have some engine of code &#8212; something resembling a factory &#8212; continuously driving it forward.</p><p>It will become similar to GitHub or CI/CD: a standard part of how serious software projects operate. I would be surprised if that did not happen.</p><h2>Start by automating the annoying parts</h2><p><strong>Latent Space:</strong> There are thousands of AI engineers at this conference. What should they do to prepare for this shift?</p><p><strong>Lloyd:</strong> Instead of only building the product directly, try building some automation toward a factory and see what it feels like.</p><p>Suppose you want an agent to implement incoming user issues automatically. What is involved in making that work? What prevents you from adopting it?</p><p>Perhaps code review is the bottleneck. Perhaps the agent is making changes, but you cannot clearly see what it did. You only discover those problems by trying to build the loop.</p><p>Get out of the mindset of building everything by hand. Find an annoying part of your job and try to create a loop that handles it for you using a factory approach.</p>]]></content:encoded></item><item><title><![CDATA[AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers]]></title><description><![CDATA[On Tuesday at the AI Engineer World's Fair, there was a lot of talk about loops, agent engineering, and the emergence of software factories. Also a hot topic: open models.]]></description><link>https://www.latent.space/p/aiewf-daily-dispatch-loops</link><guid isPermaLink="false">https://www.latent.space/p/aiewf-daily-dispatch-loops</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Wed, 01 Jul 2026 04:46:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!i4cw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i4cw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i4cw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i4cw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:810379,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i4cw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!i4cw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee341b9f-8fd3-47e5-87db-9b3d0fc72de5_1280x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Agents are here to serve you in the software factory.</figcaption></figure></div><p>Loops, loops and more loops. That word, loop, dominated conversations on day 2 of the <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer World&#8217;s Fair</a> &#8212; the first full day of keynotes and sessions. Perhaps knowing in advance what everyone would be talking about, AIEWF cofounder swyx titled his opening talk, &#8220;Loopcraft: The Art of Stacking Loops.&#8221;</p><p>swyx began by commenting on the evolution of AI engineering from 2022: from chat, to tools, to goals. &#8220;These days, we&#8217;re all about automations,&#8221; he added. &#8220;We&#8217;re all about cron jobs and loops.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!O16d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!O16d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 424w, https://substackcdn.com/image/fetch/$s_!O16d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 848w, https://substackcdn.com/image/fetch/$s_!O16d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!O16d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!O16d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg" width="1280" height="769" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:769,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:826241,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!O16d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 424w, https://substackcdn.com/image/fetch/$s_!O16d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 848w, https://substackcdn.com/image/fetch/$s_!O16d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!O16d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e4c714-6ee4-41d1-899f-e0e3afee363b_1280x769.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Allie Howe, a member of technical staff for Keycard, then introduced the main stage track for the day: Software Factories. She referenced Geoffrey Huntley&#8217;s influential article, &#8220;<a href="https://ghuntley.com/loop/">everything is a ralph loop</a>,&#8221; a theory about turning an AI coding agent into a persistent worker by repeatedly restarting it against the same spec.</p><p>Pablo Castro from Microsoft then talked about Foundry, the company&#8217;s &#8220;AI app and agent factory.&#8221; He claimed that a &#8220;learning loop&#8221; occurs when people and agents work together.</p><p>OpenAI&#8217;s Alexander Embiricos and Romain Huet were next on, and they focused a lot on Codex, the company&#8217;s coding agent. One point they made was that using multiple agents via loops can result in enhanced productivity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dfdM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dfdM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dfdM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg" width="1280" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:803731,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dfdM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dfdM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d35a5b-4958-4c34-926a-00156642dab9_1280x850.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&#8220;There will be a lot of talk today about loops,&#8221; Embiricos said. &#8220;And if you can connect the agent to not only the work that you have to do, but <em>why</em> it has to be done, that&#8217;s how you can get the agent to start to begin much more work. And then if you can connect it to what you do afterwards, review and deploy, that&#8217;s how you help it land much more work.&#8221;</p><p>This segued to a presentation by Peter Steinberger, the &#8220;ClawFather&#8221; of OpenClaw, now working for OpenAI. He too was all-in on loops, noting that he designs loops to manage agents. He added that deciding what to pay attention to is his main challenge nowadays &#8212; and that the future is &#8220;better loops&#8221; to help solve this issue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6tmE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6tmE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6tmE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg" width="1280" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:774823,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6tmE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6tmE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5478973-1c9b-43a6-a1bf-e7346be0f717_1280x850.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Software factories</h2><p>All this talk of looping led naturally to the concept of &#8220;software factories,&#8221; the subject of a presentation by Tereza T&#237;&#382;kov&#225; from a company called Factory. She defined a software factory as &#8220;the whole loop, the whole lifecycle of developing software with autonomy.&#8221; She added that this doesn&#8217;t mean just coding, but also &#8220;collecting all the signals, reacting to user feedback [and] to logs, prioritizing what&#8217;s important, then orchestrating it all.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PgI3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PgI3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PgI3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg" width="1280" height="723" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:723,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:272217,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PgI3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PgI3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366dac2b-5c86-4c55-aa69-327233d5d93a_1280x723.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Zach Lloyd from Warp also spoke about software factories; in fact, his thesis was that &#8220;software engineering will become factory engineering.&#8221; Loops in Lloyd&#8217;s framing were about improving the system.</p><p>In both T&#237;&#382;kov&#225; and Lloyd&#8217;s talks, the emphasis was on having the agents doing the building for you. &#8220;You&#8217;ll be building the thing that builds the product,&#8221; was how Lloyd put it.</p><p>Afterwards, I went down to Warp&#8217;s booth in the AIEWF expo hall and spoke to Lloyd about software factories. I particularly wanted to know why Warp, which began as a CLI tool for developers, has pivoted into a &#8216;software factory&#8217; platform where developers aren&#8217;t supposed to do coding anymore.</p><p>&#8220;The way to think of the factory is, like, pick your repos, pick the parts of the lifecycle that you want to automate, pick the ways in which you want humans to be brought into the loop,&#8221; Lloyd told me. &#8220;And different organizations [and] code bases will have different preferences for, like, do you fully automate code review [or] do you have humans do hard coding, stuff like that.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fuIy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fuIy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fuIy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg" width="1280" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:637354,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fuIy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!fuIy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5ffb36-1b2a-49c5-acaa-73af9ffe3a47_1280x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I noted that the term &#8220;factory&#8221; might be offputting to many developers, since it implies mechanized rote work &#8212; much different from the creative era of coding we&#8217;ve just come from. Lloyd recognizes this is a challenge, but he argues software factories will become a new discipline of engineering &#8212; and that it still requires problem solving.</p><p>&#8220;For better or worse, the power of these systems is so great and the ability to accelerate is so strong that just writing stuff by hand...I don&#8217;t think it&#8217;s going to make sense for very much longer,&#8221; he said.</p><p>(For more from Zach Lloyd on software factories, stay tuned for a Latent Space interview to publish shortly.)</p><h2>Forward Deployed Engineers</h2><p>Related to loops and software factories, another theme from AIEWF today was the trendy new role of Forward Deployed Engineers. In <a href="https://www.latent.space/p/forward-deployed-engineers-aiewf">an interview with Natalie Meurer</a>, Head of Agent Engineering at Sierra, I established that FDEs are also sometimes called &#8220;agent engineers.&#8221; The main point is to help organizations adapt to agents, from a development perspective.</p><p>Meurer pointed out that a lot of the work of integrating AI into companies these days is in orchestrating agents.</p><p>&#8220;In practice, most customer-specific work takes place at the orchestration layer rather than in the models themselves,&#8221; she told me.</p><p>Cursor&#8217;s VP of Forward Deployed Engineering, Pauline Brunet, also ran a session today at AIEWF, in which she positioned FDE as part of the shift to software factories. &#8220;We partner with your organization to co-design and co-build your AI software factory,&#8221; she said. &#8220;We transform how you design, develop, and maintain software across your entire life cycle.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6GAG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6GAG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6GAG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg" width="1280" height="960" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:960,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:662727,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6GAG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6GAG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6bbb3-a932-4ba4-917b-56bffdbdc915_1280x960.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>(More insights from Brunet coming in an upcoming Q&amp;A.)</p><h2>Open Source AI</h2><p>Another key theme from AIEWF today was the rise of open source AI. Zixuan Li, the head of intriguing new Chinese company Z.ai, was due to make an appearance at the conference. Because of travel issues, he couldn&#8217;t make it in person. He did make a virtual presentation, though, focusing on the company&#8217;s groundbreaking open LLM, GLM-5.2 &#8212; its &#8220;flagship model for long-horizon tasks.&#8221;</p><p>He also introduced ZCode, a harness that &#8220;supports all frontier models.&#8221; Li compared it specifically to OpenAI&#8217;s Codex.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rJ6b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rJ6b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rJ6b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg" width="1280" height="750" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:750,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:537956,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204384909?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rJ6b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rJ6b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e4d1b2-dcb8-4904-b162-65f424c8a967_1280x750.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>HuggingFace&#8217;s Thomas Wolf then interviewed Olive Song from Chinese company MiniMax, which recently released its latest open-weight model, M3.</p><p>Open source AI is a big reason why <a href="https://www.latent.space/p/ahmad-osman-local-ai">local AI is becoming more popular</a>. Ahmad Osman is the founder of Osmantic, a company building open source software for deploying and operating local AI systems. He spoke to us today and noted that open models have improved dramatically in recent times.</p><p>&#8220;Architectures are becoming more efficient, and many small improvements compound,&#8221; he said. &#8220;Once a frontier lab demonstrates that a capability is possible, the open source ecosystem can work backwards from that and find ways to reproduce it more efficiently.&#8221;</p><h2>Conclusion</h2><p>Those were the big trends from day 2 of the AI Engineer World&#8217;s Fair. I&#8217;ll be back tomorrow with all the action and analysis from day 3. Don&#8217;t forget to <a href="https://www.youtube.com/@aiDotEngineer/streams">tune into the keynotes</a> on YouTube if you&#8217;re following from work or home.</p>]]></content:encoded></item><item><title><![CDATA[[AINews] Sonnet 5 today, and Fable 5 tomorrow]]></title><description><![CDATA[Everything is open again!]]></description><link>https://www.latent.space/p/ainews-sonnet-5-today-and-fable-5</link><guid isPermaLink="false">https://www.latent.space/p/ainews-sonnet-5-today-and-fable-5</guid><pubDate>Wed, 01 Jul 2026 03:01:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!V4wu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMF3K5vakAAfUEM.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In separate announcements, <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a> was released today, and <a href="https://x.com/anthropicai/status/2072106151890809341?s=46">Fable/Mythos 5 were approved</a> to be released again after some work with the government. The <a href="https://x.com/theo/status/2072068395529576912">primary discussion around Sonnet 5&#8217;s efficiency</a> was a damper on the excitement, driven by <a href="https://x.com/simonw/status/2072068898648949184">tokenizer changes</a> and <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">3-6x more turn taking</a> in benchmarks:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/theo/status/2072068395529576912&quot;,&quot;full_text&quot;:&quot;Oh my god, Sonnet 5 was MORE EXPENSIVE THAN FABLE to run the whole bench &#128128;&quot;,&quot;username&quot;:&quot;theo&quot;,&quot;name&quot;:&quot;Theo - t3.gg&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1909353910130950147/EeSGdgA5_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-30T21:22:57.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HMF3K5vakAAfUEM.png&quot;,&quot;link_url&quot;:&quot;https://t.co/ZGNsWsVIay&quot;}],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;Sonnet 5 cost MORE than Opus 4.8 on the Artificial Analysis Intelligence Index&quot;,&quot;username&quot;:&quot;theo&quot;,&quot;name&quot;:&quot;Theo - t3.gg&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1909353910130950147/EeSGdgA5_normal.jpg&quot;},&quot;reply_count&quot;:144,&quot;retweet_count&quot;:86,&quot;like_count&quot;:2137,&quot;impression_count&quot;:171344,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>Our newest staff writer <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Richard MacManus&quot;,&quot;id&quot;:232063,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c4ca3255-4ccf-497e-a04f-219d65fba554_2048x2048.png&quot;,&quot;uuid&quot;:&quot;6aabcc91-4722-42e7-9dac-63c3acb483c7&quot;}" data-component-name="MentionToDOM"></span>  is reporting on the ground from AIE, and you can catch swyx and other keynote speakers on the stream today:</p><div id="youtube2-htM02KMNZnk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;htM02KMNZnk&quot;,&quot;startTime&quot;:&quot;16s&quot;,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/htM02KMNZnk?start=16s&amp;rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><blockquote><p>AI News for 6/29/2026-6/30/2026. We checked 12 subreddits, <a href="https://twitter.com/i/lists/1585430245762441216">544 Twitters</a> and no further Discords. <a href="https://news.smol.ai/">AINews&#8217; website</a> lets you search all past issues. As a reminder, <a href="https://www.latent.space/p/2026">AINews is now a section of Latent Space</a>. You can <a href="https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack">opt in/out</a> of email frequencies!</p></blockquote><div><hr></div><h1><strong>AI Twitter Recap</strong></h1><p><strong>Anthropic launched Claude Sonnet 5 as its new default mid-tier frontier model, with immediate rollout across Claude, Claude Code, API, and ecosystem partners.</strong></p><ul><li><p>Anthropic officially announced <strong>Claude Sonnet 5</strong> as &#8220;our most agentic Sonnet yet,&#8221; emphasizing planning, browser/terminal tool use, and autonomous execution that previously &#8220;required larger and more expensive models&#8221; (<a href="https://x.com/claudeai/status/2072017450611142835">@claudeai</a>)</p></li><li><p>Anthropic&#8217;s developer account said Sonnet 5 offers <strong>top-tier coding and tool-use performance at Sonnet pricing</strong>, with a <strong>1M-token context window</strong>, and is the <strong>new default in Claude Code for Pro users</strong> and available on the Claude Platform including <strong>API and Managed Agents</strong> (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>)</p></li><li><p>Anthropic kept the standard list price at <strong>$3/M input tokens and $15/M output tokens</strong>, but introduced a <strong>promotional rate of $2/M input and $10/M output through Aug. 31 / Sept. 1 depending on the post</strong> (<a href="https://x.com/kimmonismus/status/2072019015577333804">@kimmonismus</a>, <a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p>Sonnet 5 surfaced first through leaks and client-side sightings: leakers claimed <strong>knowledge cutoff January 2026</strong>, <strong>$2/$10 promo pricing</strong>, and a <strong>1M-context variant</strong> before launch (<a href="https://x.com/kimmonismus/status/2071953298169778636">@kimmonismus</a>); users then reported it appearing in the <strong>model selector</strong>, <strong>Claude Code 2.1.197</strong>, <strong>Anthropic GitHub</strong>, and finally going live in accounts including <strong>Germany</strong> (<a href="https://x.com/kimmonismus/status/2071971743556628668">@kimmonismus</a>, <a href="https://x.com/scaling01/status/2071969195726659829">@scaling01</a>, <a href="https://x.com/scaling01/status/2072014332104265884">@scaling01</a>, <a href="https://x.com/kimmonismus/status/2072017872478470586">@kimmonismus</a>)</p></li><li><p>Anthropic simultaneously expanded platform support around the launch: <strong>Claude Desktop on Linux (Ubuntu/Debian beta)</strong> with Claude Code/Cowork/chat on paid plans, though <strong>Computer Use was not included</strong> in that Linux release (<a href="https://x.com/ClaudeDevs/status/2071988881717871065">@ClaudeDevs</a>, <a href="https://x.com/ClaudeDevs/status/2071988883802444125">@ClaudeDevs</a>)</p></li><li><p>Anthropic also shipped <strong>Managed Agents</strong> updates&#8212;streaming session deltas, per-session overrides, webhook events, reverse pagination, credential injection scoping, and an observability tab with token/tool metrics&#8212;making the release as much platform/integration story as raw model story (<a href="https://x.com/ClaudeDevs/status/2072058428424589412">@ClaudeDevs</a>, <a href="https://x.com/ClaudeDevs/status/2072058433097122145">@ClaudeDevs</a>)</p></li></ul><h2><strong>Launch timeline and pre-release narrative</strong></h2><p>The launch was preceded by a large rumor cycle centered on <strong>Sonnet 5 + Fable 5</strong>.</p><ul><li><p>Earlier app-string sleuthing suggested Anthropic was preparing to put <strong>&#8220;Fable 5&#8221; behind a separate usage-credit system billed outside existing plans</strong>, with <strong>identity verification</strong> language appearing nearby; that fed speculation that access would be gated and more regulated than existing plans (<a href="https://x.com/kimmonismus/status/2071868011804266828">@kimmonismus</a>)</p></li><li><p>This triggered concern that Sonnet 5 might launch as the <strong>widely accessible but weaker</strong> companion to a stronger, more restricted <strong>Fable 5</strong>, possibly with regional access issues, especially in Europe (<a href="https://x.com/kimmonismus/status/2071899142616408377">@kimmonismus</a>)</p></li><li><p>Additional rumor posts tied a potential Sonnet 5 release directly to a <strong>Fable 5 re-release</strong>, with some users explicitly saying they assumed Sonnet 5 would &#8220;at least&#8221; come with Fable news (<a href="https://x.com/kimmonismus/status/2071941904636531167">@kimmonismus</a>, <a href="https://x.com/kimmonismus/status/2071953298169778636">@kimmonismus</a>)</p></li><li><p>After launch, that expectation went unmet. Multiple reactions framed the absence of Fable 5 as the real story: &#8220;instead we got sonnet 5&#8221; (<a href="https://x.com/kimmonismus/status/2072058904352002271">@kimmonismus</a>) and &#8220;It&#8217;s been 18 days since Fable 5 was banned&#8221; (<a href="https://x.com/theo/status/2072058513669693608">@theo</a>)</p></li></ul><h2><strong>Official positioning vs independent interpretation</strong></h2><h3><strong>Official/vendor framing</strong></h3><p>Anthropic and downstream partners framed Sonnet 5 around <strong>agentic capability, coding, tool use, and cost-performance</strong>.</p><ul><li><p>Official claim: Sonnet 5 is the <strong>&#8220;most agentic Sonnet yet&#8221;</strong> and can make plans, use browsers/terminals, and operate autonomously at a level that recently required larger models (<a href="https://x.com/claudeai/status/2072017450611142835">@claudeai</a>)</p></li><li><p>Anthropic&#8217;s dev account positioned it as <strong>frontier-quality coding and tool use at Sonnet pricing</strong>, explicitly highlighting <strong>1M context</strong> and broad platform availability (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>)</p></li><li><p>Anthropic-linked summary posts stressed that Sonnet 5 is <strong>safer than Sonnet 4.6 overall</strong>, with lower <strong>hallucination</strong> and <strong>sycophancy</strong>, and that <strong>cyber safeguards are on by default</strong>, while still acknowledging <strong>Opus remains stronger for serious cyber work</strong> (<a href="https://x.com/kimmonismus/status/2072019015577333804">@kimmonismus</a>)</p></li><li><p>Anthropic also provided migration tooling/documentation, saying the <strong>claude-api skill</strong> helps tune prompts, recommend effort levels, and configure advisor mode for Sonnet 5 (<a href="https://x.com/ClaudeDevs/status/2072018517898272844">@ClaudeDevs</a>)</p></li></ul><h3><strong>Independent/third-party evaluation framing</strong></h3><p>Third parties largely agreed Sonnet 5 is a <strong>real improvement over Sonnet 4.6</strong>, but disputed whether it merits a &#8220;5.0&#8221; naming step or its effective price/performance relative to Opus and peers.</p><ul><li><p>Cursor said Sonnet 5 is a <strong>meaningful step up</strong> on <strong>CursorBench: 57% vs 49%</strong> for Sonnet 4.6 (<a href="https://x.com/cursor_ai/status/2072020786181988418">@cursor_ai</a>)</p></li><li><p>Cognition said Sonnet 5 <strong>outperforms Opus 4.8 on FrontierCode Extended</strong>, posting <strong>53.8% score</strong> and <strong>57.6% pass rate</strong>, while noting benchmark rankings may shift slightly after upcoming adjustments (<a href="https://x.com/cognition/status/2072022778144821292">@cognition</a>, <a href="https://x.com/cognition/status/2072022781043028182">@cognition</a>)</p></li><li><p>Cline highlighted <strong>Opus 4.8-level performance on Terminal-Bench for less than half the cost</strong>, plus improved resistance to <strong>prompt-injection hijacks</strong> for &#8220;--yolo coders&#8221; (<a href="https://x.com/cline/status/2072051144436928727">@cline</a>)</p></li><li><p>FactoryAI, Perplexity, Cursor, Devin, Droid, Agent Arena, and VS Code all quickly added support or availability announcements, indicating the ecosystem saw it as a relevant default model even where user enthusiasm was mixed (<a href="https://x.com/FactoryAI/status/2072021755619864778">@FactoryAI</a>, <a href="https://x.com/perplexity_ai/status/2072030042994160028">@perplexity_ai</a>, <a href="https://x.com/AravSrinivas/status/2072031649693675810">@AravSrinivas</a>, <a href="https://x.com/code/status/2072029026881859987">@code</a>, <a href="https://x.com/arena/status/2072035566829568111">@arena</a>, <a href="https://x.com/cognition/status/2072022778144821292">@cognition</a>)</p></li></ul><h2><strong>Technical details</strong></h2><h3><strong>Core product specs and pricing</strong></h3><ul><li><p><strong>Context window:</strong> <strong>1 million tokens</strong> (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Standard pricing:</strong> <strong>$3/M input, $15/M output</strong> (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Promotional pricing:</strong> <strong>$2/M input, $10/M output</strong> until <strong>Aug. 31 / Sept. 1</strong> depending on wording of the post (<a href="https://x.com/kimmonismus/status/2072019015577333804">@kimmonismus</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Cache pricing:</strong> <strong>25% premium for cache writes ($3.75/M)</strong>, <strong>90% discount for cache hits ($0.3/M)</strong>, <strong>5-minute TTL</strong> (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Effort settings:</strong> Sonnet 5 adds <strong>xhigh</strong>, for <strong>5 effort levels total</strong> matching Opus 4.8: <strong>max, xhigh, high, medium, low</strong> (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Knowledge cutoff (rumored pre-launch):</strong> <strong>January 2026</strong> (<a href="https://x.com/kimmonismus/status/2071953298169778636">@kimmonismus</a>)</p></li></ul><h3><strong>Benchmarks and measured deltas</strong></h3><p>A key part of the discussion was that Sonnet 5 improved substantially over 4.6, but usually <strong>did not exceed Opus 4.8 on broad intelligence aggregates</strong>.</p><ul><li><p><strong>CursorBench:</strong> <strong>57%</strong> for Sonnet 5 vs <strong>49%</strong> for Sonnet 4.6 (<a href="https://x.com/cursor_ai/status/2072020786181988418">@cursor_ai</a>)</p></li><li><p><strong>Artificial Analysis Intelligence Index:</strong> Sonnet 5 scores <strong>53</strong>, a <strong>+6</strong> over Sonnet 4.6, placing it <strong>#5 overall</strong>, roughly tied with <strong>GPT-5.5 high reasoning</strong>, but still behind <strong>Opus 4.7/4.8</strong> (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Artificial Analysis token usage:</strong> Sonnet 5 used <strong>~69k output tokens per task on average</strong>, about <strong>40% more output tokens</strong> than Sonnet 4.6 (<a href="https://x.com/ArtificialAnlys/status/2072062598187765893">@ArtificialAnlys</a>)</p></li><li><p><strong>Artificial Analysis task cost:</strong> at standard pricing, Sonnet 5 cost <strong>$2.29 per Intelligence Index task</strong>, about <strong>2x Sonnet 4.6</strong> and <strong>~15% more than Opus 4.8</strong>, despite lower per-token price, because of higher token usage (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>Agentic turns:</strong> Sonnet 5 used <strong>~3x the agentic turns</strong> of Sonnet 4.6 on <strong>AA-Briefcase</strong> and <strong>GDPval-AA</strong>, and <strong>max effort</strong> used around <strong>6x more turns</strong> than <strong>low effort</strong> on GDPval-AA (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p><strong>CritPt frontier physics benchmark:</strong> Sonnet 5 scored <strong>17%</strong>, <strong>+14 points</strong> over its predecessor, but still behind <strong>GLM-5.2</strong>, <strong>Claude Opus</strong>, <strong>Fable</strong>, and <strong>GPT-5.5</strong> variants (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p>Artificial Analysis also reported notable improvements over Sonnet 4.6 on <strong>Terminal-Bench v2.1 (+9)</strong>, <strong>Humanity&#8217;s Last Exam (+10)</strong>, and <strong>SciCode (+7)</strong> (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p>Cognition&#8217;s <strong>FrontierCode Extended</strong> result: <strong>53.8% score</strong>, <strong>57.6% pass rate</strong>, ahead of Opus 4.8 in their current evaluation (<a href="https://x.com/cognition/status/2072022781043028182">@cognition</a>)</p></li><li><p>Max Bittker noted <strong>Runescape benchmark</strong> scores improved a lot over Sonnet 4.6, but were still behind nearby Pareto competitors such as <strong>GLM 5.2</strong> and <strong>Gemini 3.5 Flash</strong> (<a href="https://x.com/maxbittker/status/2072054926746779806">@maxbittker</a>)</p></li></ul><h3><strong>Tokenization and effective cost quirks</strong></h3><p>One underappreciated technical detail was the tokenizer/effective billing behavior.</p><ul><li><p>Simon Willison noted the <strong>new tokenizer</strong> makes Sonnet 5 <strong>~1.4x more expensive for English</strong>, <strong>~1.33x for Spanish</strong>, and <strong>roughly the same for Simplified Mandarin</strong> (<a href="https://x.com/simonw/status/2072068898648949184">@simonw</a>)</p></li><li><p>This matters because many users compared only list prices, while evaluators and power users focused on <strong>cost per solved task</strong>, not just <strong>cost per token</strong></p></li></ul><h2><strong>Facts vs opinions</strong></h2><h3><strong>Factual claims supported by official or benchmark posts</strong></h3><ul><li><p>Sonnet 5 launched officially and is available in <strong>Claude, Claude Code, API, Managed Agents</strong>, and many partner products (<a href="https://x.com/claudeai/status/2072017450611142835">@claudeai</a>, <a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>)</p></li><li><p>It has a <strong>1M-token context window</strong> (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>)</p></li><li><p>Standard pricing is <strong>$3/$15 per million input/output tokens</strong> with a temporary promo of <strong>$2/$10</strong> (<a href="https://x.com/ClaudeDevs/status/2072018504392601762">@ClaudeDevs</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p>Third-party results show meaningful gains over Sonnet 4.6 on coding/agentic benchmarks including CursorBench, FrontierCode Extended, and Artificial Analysis (<a href="https://x.com/cursor_ai/status/2072020786181988418">@cursor_ai</a>, <a href="https://x.com/cognition/status/2072022781043028182">@cognition</a>, <a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li><li><p>Artificial Analysis found Sonnet 5 can cost <strong>more per task than Opus 4.8</strong> because it uses more tokens/turns (<a href="https://x.com/ArtificialAnlys/status/2072062592923930666">@ArtificialAnlys</a>)</p></li></ul><h3><strong>Rumors / unverified claims</strong></h3><ul><li><p><strong>Fable 5</strong> billing changes, identity verification, and regulatory linkage came from app-string interpretation and user speculation, not from an official launch note (<a href="https://x.com/kimmonismus/status/2071868011804266828">@kimmonismus</a>)</p></li><li><p><strong>January 2026 knowledge cutoff</strong> and some launch/pricing details were leaked before confirmation (<a href="https://x.com/kimmonismus/status/2071953298169778636">@kimmonismus</a>)</p></li><li><p>Claims that Sonnet 5 was <strong>intentionally nerfed</strong>, <strong>self-distilled just enough to remain below Opus</strong>, or launched due to a <strong>soft ban on frontier capabilities</strong> are opinions/speculation, not evidenced in the official materials (<a href="https://x.com/scaling01/status/2072039834529435674">@scaling01</a>, <a href="https://x.com/z4y5f3/status/2072028918622622026">@z4y5f3</a>, <a href="https://x.com/kimmonismus/status/2072027861385466123">@kimmonismus</a>)</p></li></ul><h3><strong>Interpretive opinions</strong></h3><ul><li><p>Positive interpretation: Sonnet 5 is the kind of <strong>smaller/cheaper model improvement</strong> that matters most for <strong>parallel workflows, long-running agents, and production coding systems</strong> (<a href="https://x.com/The_Whole_Daisy/status/2072019554935652746">@The_Whole_Daisy</a>, <a href="https://x.com/omarsar0/status/2072022542521438300">@omarsar0</a>, <a href="https://x.com/skirano/status/2072044693798412782">@skirano</a>)</p></li><li><p>Negative interpretation: Sonnet 5 is <strong>underwhelming</strong>, overpriced in practice, and mislabeled as &#8220;5&#8221; when its aggregate capability looks closer to <strong>4.8/4.9</strong> than a major generational leap (<a href="https://x.com/kimmonismus/status/2072027861385466123">@kimmonismus</a>, <a href="https://x.com/scaling01/status/2072039834529435674">@scaling01</a>, <a href="https://x.com/DeryaTR_/status/2072051617298293199">@DeryaTR_</a>)</p></li><li><p>Neutral/engineering interpretation: This is a <strong>production-friendly release</strong> more than a hype release&#8212;better on coding/agents, broadly deployable, but not a flagship-redefining jump (<a href="https://x.com/dejavucoder/status/2072020732226478192">@dejavucoder</a>, <a href="https://x.com/OpenAIDevs/status/2072036305442406772">@OpenAIDevs</a>)</p></li></ul><h2><strong>Different opinions</strong></h2><h3><strong>Supporting views</strong></h3><ul><li><p><strong>Production users benefit most.</strong> Several posters argued Sonnet 5 is exactly the kind of model teams want for <strong>long-running agents</strong>, <strong>coding loops</strong>, and <strong>tool-use reliability</strong>, even if it doesn&#8217;t win every static benchmark (<a href="https://x.com/omarsar0/status/2072022542521438300">@omarsar0</a>, <a href="https://x.com/skirano/status/2072044693798412782">@skirano</a>)</p></li><li><p><strong>Smaller-model launches matter.</strong> Power users can underappreciate how much value comes from making a cheaper/default-tier model stronger, because that unlocks more parallel agents and redundancy in workflows (<a href="https://x.com/The_Whole_Daisy/status/2072019554935652746">@The_Whole_Daisy</a>)</p></li><li><p><strong>Coding benchmarks are strong.</strong> Cursor and Cognition both posted substantial results in practical coding/evaluation harnesses (<a href="https://x.com/cursor_ai/status/2072020786181988418">@cursor_ai</a>, <a href="https://x.com/cognition/status/2072022781043028182">@cognition</a>)</p></li><li><p><strong>Security angle improved.</strong> Cline highlighted better resistance to prompt-injection/hijack attempts, relevant to autonomous terminal/browser usage (<a href="https://x.com/cline/status/2072051144436928727">@cline</a>)</p></li></ul><h3><strong>Critical views</strong></h3><p>The strongest criticism focused on <strong>naming, absent Fable 5, and poor task-level cost efficiency</strong>.</p><ul><li><p><strong>Naming criticism:</strong> users argued &#8220;Sonnet 5&#8221; implies a major-version leap, while evals suggest something closer to <strong>Sonnet 4.8/4.9</strong> (<a href="https://x.com/kimmonismus/status/2072027861385466123">@kimmonismus</a>, <a href="https://x.com/teortaxesTex/status/2072021520352772185">@teortaxesTex</a>)</p></li><li><p><strong>Benchmark criticism:</strong> multiple users stressed Sonnet 5 still trails <strong>Opus 4.8</strong> &#8220;across all evals&#8221; or on broad intelligence measures (<a href="https://x.com/kimmonismus/status/2072027861385466123">@kimmonismus</a>, <a href="https://x.com/theo/status/2072066764465393917">@theo</a>)</p></li><li><p><strong>Cost-per-task criticism:</strong> this became the most technically grounded negative theme. Theo, Yuchen Jin, Scaling01, and Kimmonismus all amplified that Sonnet 5 can be <strong>more expensive than Opus 4.8 or even Fable on actual evaluated tasks</strong> due to verbosity/turn count (<a href="https://x.com/theo/status/2072066764465393917">@theo</a>, <a href="https://x.com/theo/status/2072068395529576912">@theo</a>, <a href="https://x.com/Yuchenj_UW/status/2072070274300948497">@Yuchenj_UW</a>, <a href="https://x.com/kimmonismus/status/2072072593109315855">@kimmonismus</a>, <a href="https://x.com/scaling01/status/2072071305281540338">@scaling01</a>)</p></li><li><p><strong>Launch disappointment tied to Fable 5:</strong> critics saw Sonnet 5 as a consolation release while the real frontier model remained withheld or constrained (<a href="https://x.com/kimmonismus/status/2072027861385466123">@kimmonismus</a>, <a href="https://x.com/theo/status/2072058513669693608">@theo</a>, <a href="https://x.com/scaling01/status/2072044421634281636">@scaling01</a>)</p></li></ul><h3><strong>Neutral / mixed takes</strong></h3><ul><li><p><strong>&#8220;Production people will be happy; personal wow-factor is low.&#8221;</strong> That succinctly captures a recurring mixed reaction (<a href="https://x.com/dejavucoder/status/2072020732226478192">@dejavucoder</a>)</p></li><li><p><strong>Good release, bad expectation management.</strong> Some users seemed less upset by the model itself than by the implication that a &#8220;5.0&#8221; label and rumor cycle primed people for a more dramatic frontier jump</p></li><li><p><strong>Agentic quality may be undermeasured.</strong> Some believed traditional benchmark comparisons may underrate improvements in what one poster called the model&#8217;s <strong>&#8220;working mind&#8221;</strong> on long-horizon tasks (<a href="https://x.com/skirano/status/2072044693798412782">@skirano</a>)</p></li></ul><h2><strong>Ecosystem rollout</strong></h2><p>Sonnet 5 was adopted unusually quickly across the coding-agent ecosystem, which is itself evidence of where the market thinks the value lies.</p><ul><li><p><strong>Cursor</strong> added Sonnet 5 and published CursorBench deltas (<a href="https://x.com/cursor_ai/status/2072020786181988418">@cursor_ai</a>)</p></li><li><p><strong>Devin Desktop / CLI</strong> added it and claimed FrontierCode Extended outperformance versus Opus 4.8, plus temporary <strong>~30% lower quota usage than Sonnet 4.6</strong> through Aug. 31 (<a href="https://x.com/cognition/status/2072022778144821292">@cognition</a>, <a href="https://x.com/cognition/status/2072022784084000810">@cognition</a>)</p></li><li><p><strong>Cline</strong> added support and emphasized Terminal-Bench/cyber-hijack robustness (<a href="https://x.com/cline/status/2072051144436928727">@cline</a>)</p></li><li><p><strong>FactoryAI Droid</strong> added Sonnet 5 at <strong>1/3 off until Aug. 31</strong> (<a href="https://x.com/FactoryAI/status/2072021755619864778">@FactoryAI</a>)</p></li><li><p><strong>Perplexity</strong> added Sonnet 5 for Pro/Max and as a <strong>Computer orchestrator model</strong> (<a href="https://x.com/perplexity_ai/status/2072030042994160028">@perplexity_ai</a>, <a href="https://x.com/AravSrinivas/status/2072031649693675810">@AravSrinivas</a>)</p></li><li><p><strong>VS Code / @code</strong> rolled it out (<a href="https://x.com/code/status/2072029026881859987">@code</a>)</p></li><li><p><strong>Arena</strong> added Sonnet 5 to Agent Arena and other arenas (<a href="https://x.com/arena/status/2072035566829568111">@arena</a>)</p></li></ul><p>This rollout pattern reinforces that Sonnet 5 is being treated less as a chatbot headline and more as a <strong>default workhorse model for agentic software stacks</strong>.</p><h2><strong>Context</strong></h2><p>Sonnet has historically been Anthropic&#8217;s <strong>price/performance workhorse</strong> and the model most likely to be used at scale in products like coding assistants, managed agents, and enterprise automation. That context matters for why the discourse split:</p><ul><li><p>Frontier-watchers expected a <strong>headline &#8220;5.x&#8221; event</strong></p></li><li><p>Builders wanted a <strong>better reliable default model</strong></p></li><li><p>Power users benchmarked <strong>per solved task</strong>, not <strong>per token</strong></p></li><li><p>Policy-aware observers interpreted the absence of <strong>Fable 5</strong> and the earlier <strong>ID-verification/credit rumors</strong> as signs of tightening governance or staged access</p></li></ul><p>The launch also lands in a market where model differentiation is increasingly about:</p><ul><li><p><strong>long-horizon tool use</strong></p></li><li><p><strong>agent reliability</strong></p></li><li><p><strong>token efficiency</strong></p></li><li><p><strong>effective cost per completed task</strong></p></li><li><p><strong>integration into work environments</strong> rather than pure chat demos</p></li></ul><p>That is why reactions ranged from &#8220;clear upgrade&#8221; to &#8220;worst Anthropic launch.&#8221; Both are responding to real but different axes:</p><ul><li><p>On <strong>absolute capability vs Sonnet 4.6</strong>, it looks materially better</p></li><li><p>On <strong>headline frontier progress vs Opus/Fable expectations</strong>, it disappointed many</p></li><li><p>On <strong>list price</strong>, it looks affordable</p></li><li><p>On <strong>task-level cost</strong>, it can look surprisingly expensive</p></li><li><p>On <strong>ecosystem utility</strong>, it was immediately embraced</p></li></ul><p><strong>China models, infrastructure, and open-weight competition</strong></p><ul><li><p>Meituan&#8217;s release drew the most attention outside Sonnet: an <strong>open-weights 1.6T-parameter model</strong> from a major Chinese delivery company, with discussion centering on how non-obvious Chinese incumbents can fund serious frontier-scale efforts (<a href="https://x.com/JosephJacks_/status/2071858781521342568">@JosephJacks_</a>, <a href="https://x.com/natolambert/status/2071972882264268923">@natolambert</a>, <a href="https://x.com/teortaxesTex/status/2071906284958294419">@teortaxesTex</a>)</p></li><li><p>Technical scrutiny focused on hardware and scale details: claims that Meituan used <strong>CloudMatrix 384 pods in &#8220;910B mode&#8221;</strong>, implying <strong>~25K chips not 50K GPUs-equivalent</strong>, while critics compared that to a future <strong>Huawei 950DT SuperPod with 8192 chips</strong> possibly outperforming the whole setup (<a href="https://x.com/teortaxesTex/status/2071888424823325139">@teortaxesTex</a>, <a href="https://x.com/teortaxesTex/status/2071889274954260720">@teortaxesTex</a>)</p></li><li><p>DSpark/DeepSeek infra remained a major subtheme: posters highlighted <strong>TPOT of 2.9&#8211;5.2 ms</strong>, possible <strong>50% throughput</strong> gains or <strong>60% interactivity</strong> gains across Chinese providers, and the view that DeepSeek&#8217;s infra open-sourcing is creating broad economic spillovers (<a href="https://x.com/teortaxesTex/status/2071879186373923284">@teortaxesTex</a>, <a href="https://x.com/teortaxesTex/status/2071873225881989424">@teortaxesTex</a>, <a href="https://x.com/Xianbao_QIAN/status/2071917185380073611">@Xianbao_QIAN</a>)</p></li><li><p>Huawei/Pangu and broader domestic stack momentum also came up: <strong>Pangu 92B / 6B active MoE</strong> open-sourcing in July was flagged, alongside repeated arguments that Chinese labs now have the software and architecture maturity to train near-frontier models on domestic hardware (<a href="https://x.com/teortaxesTex/status/2071890951816003663">@teortaxesTex</a>, <a href="https://x.com/teortaxesTex/status/2072038240027131963">@teortaxesTex</a>)</p></li></ul><p><strong>Inference, chips, and systems</strong></p><ul><li><p>Etched&#8217;s stealth exit dominated hardware news: the company said it has <strong>$800M raised</strong>, <strong>$1B+ customer contracts</strong>, successful <strong>A0 tapeout</strong>, early <strong>SOTA throughput/latency/power efficiency</strong> in customer tests, and first racks shipping this summer (<a href="https://x.com/Etched/status/2071972062202343590">@Etched</a>)</p></li><li><p>Follow-on commentary described two notable hardware ideas: <strong>low-voltage inference</strong> to avoid thermal throttling under sustained load, and <strong>cluster-scale memory</strong> aimed at SRAM-like access speeds with larger pooled memory for long-context / giant-model inference (<a href="https://x.com/LiorOnAI/status/2072017343262466097">@LiorOnAI</a>)</p></li><li><p>OpenAI also reportedly found an inference optimization that <strong>more than halved inference costs</strong>, reducing logged-out ChatGPT traffic to &#8220;a couple hundred&#8221; GPUs at one point; several posts noted the strategic implication for margins and API pricing rather than the unknown exact trick (<a href="https://x.com/steph_palazzolo/status/2071972245849710938">@steph_palazzolo</a>, <a href="https://x.com/kimmonismus/status/2071987406656655416">@kimmonismus</a>)</p></li><li><p>A strong technical explainer traced NVIDIA programming&#8217;s evolution from Volta to Blackwell: from synchronous thread-centric CUDA to <strong>asynchronous dataflow across Tensor Cores, memory engines, barriers, TMA/TMEM</strong>, with detailed compute/bandwidth ratios for <strong>V100, A100, H100, B100</strong> and examples from <strong>FlashAttention-3</strong> and <strong>FlashMLA</strong> (<a href="https://x.com/ZhihuFrontier/status/2071871535430926400">@ZhihuFrontier</a>)</p></li></ul><p><strong>Agents, loops, evals, and memory</strong></p><ul><li><p>AI Engineer World Fair discourse strongly converged on <strong>&#8220;loops&#8221; / &#8220;loop engineering&#8221;</strong> as the new practical frame for agentic software: Andrew Ng described <strong>agentic coding</strong>, <strong>developer feedback</strong>, and <strong>external feedback</strong> loops as the operating model for AI-native product development (<a href="https://x.com/AndrewYNg/status/2071988145667928442">@AndrewYNg</a>)</p></li><li><p>The same theme appeared across conference chatter and tools: posts noted &#8220;loopcraft&#8221; in the keynote and heavy reuse of the term by OpenAI/Microsoft speakers and Peter Steinberger (<a href="https://x.com/latentspacepod/status/2072003484120203362">@latentspacepod</a>, <a href="https://x.com/swyx/status/2071977886991679715">@swyx</a>)</p></li><li><p>Agent evaluation infrastructure also advanced: LangChain integrated <strong>Harbor</strong> with <strong>Deep Agents, LangSmith Sandboxes, and Observability</strong>, positioning reproducible environment-based evals as becoming the standard for long-running/stateful agents (<a href="https://x.com/LangChain/status/2071978566691049559">@LangChain</a>, <a href="https://x.com/hwchase17/status/2071974139926294897">@hwchase17</a>)</p></li><li><p>Memory was another recurring topic: Harrison Chase and others highlighted <strong>wiki-style memory</strong> as one of the most promising agent memory patterns, with examples including <strong>DeepWiki, AutoWiki, LLM Wiki</strong>, and repeated emphasis that the hard part is not the storage backend but the condensation/retrieval process (<a href="https://x.com/hwchase17/status/2071963841009942671">@hwchase17</a>, <a href="https://x.com/BraceSproul/status/2071982037276475502">@BraceSproul</a>)</p></li></ul><p><strong>Models, benchmarks, and media releases</strong></p><ul><li><p>Google launched two media models: <strong>Nano Banana 2 Lite</strong> for images and <strong>Gemini Omni Flash</strong> for video generation/editing. Reported specs included <strong>&lt;4s image generation</strong>, <strong>$0.034 per 1K image</strong>, and <strong>$0.10/sec</strong> for Omni Flash video, with strong early Arena placement (<a href="https://x.com/GoogleDeepMind/status/2071988044878516466">@GoogleDeepMind</a>, <a href="https://x.com/OfficialLoganK/status/2071988351083921690">@OfficialLoganK</a>, <a href="https://x.com/arena/status/2072049269054562711">@arena</a>)</p></li><li><p>Open-weight model discussions remained active: GLM-5.2 was repeatedly cited as the strongest open model on some intelligence/enterprise benchmarks, though criticized for verbosity and high output-token usage (<a href="https://x.com/ArtificialAnlys/status/2072022576394821859">@ArtificialAnlys</a>, <a href="https://x.com/RajeswarSai/status/2072006835444347390">@RajeswarSai</a>)</p></li><li><p>Microsoft reportedly released a <strong>4B GUI agent</strong> with a jump from <strong>39.8% to 82.9% task success</strong> according to one summary post, though without source detail in the tweet itself (<a href="https://x.com/HuggingPapers/status/2071951218889339131">@HuggingPapers</a>)</p></li><li><p>OpenAI introduced <strong>GeneBench-Pro</strong>, a benchmark for realistic computational biology agent work rather than biology QA, while OpenAI Devs also published a deep debugging writeup on a year-long infra crash hunt (<a href="https://x.com/OpenAI/status/2072004836674167294">@OpenAI</a>, <a href="https://x.com/OpenAIDevs/status/2071995642436800916">@OpenAIDevs</a>)</p></li></ul><p><strong>Open-source/local AI and tooling</strong></p><ul><li><p>Hugging Face added a <strong>hardware filter</strong> for model discovery, letting users filter by GPU/CPU/Apple Silicon compatibility; this was framed as making local/open models much more usable at scale (<a href="https://x.com/victormustar/status/2071930123549290707">@victormustar</a>, <a href="https://x.com/mervenoyann/status/2071941995514237193">@mervenoyann</a>, <a href="https://x.com/ClementDelangue/status/2071951499660292496">@ClementDelangue</a>)</p></li><li><p>Several posts explicitly linked local models to resilience against platform restrictions and identity verification concerns on proprietary systems (<a href="https://x.com/kimmonismus/status/2071877617150517526">@kimmonismus</a>, <a href="https://x.com/JayAlammar/status/2071950697096987040">@JayAlammar</a>)</p></li><li><p>New open benchmarks and tools included <strong>IFStruct</strong> for output validity/schema following (<a href="https://x.com/maximelabonne/status/2071959319923380481">@maximelabonne</a>), <strong>CS2-10k</strong> with <strong>600K+ egocentric gameplay videos / 10K+ hours</strong> for world models and action-conditioned generation (<a href="https://x.com/RekaAILabs/status/2071970771233038475">@RekaAILabs</a>), and <strong>Buckets S3 API</strong> for Hugging Face storage interoperability (<a href="https://x.com/vanstriendaniel/status/2071919131058712878">@vanstriendaniel</a>)</p></li><li><p>Sebastian Raschka&#8217;s <strong>Build a Reasoning Model (From Scratch)</strong> launch was one of the highest-engagement educational items: <strong>440 full-color pages</strong> on inference scaling, RL, and distillation (<a href="https://x.com/rasbt/status/2071945864088535126">@rasbt</a>)</p></li></ul><div><hr></div><h1><strong>AI Reddit Recap</strong></h1><h2><strong>/r/LocalLlama + /r/localLLM Recap</strong></h2><p></p>
      <p>
          <a href="https://www.latent.space/p/ainews-sonnet-5-today-and-fable-5">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Forward Deployed Engineers and the future of software engineering]]></title><description><![CDATA[Sierra's Natalie Meurer on why product engineers and forward deployed engineers are starting to converge.]]></description><link>https://www.latent.space/p/forward-deployed-engineers-aiewf</link><guid isPermaLink="false">https://www.latent.space/p/forward-deployed-engineers-aiewf</guid><dc:creator><![CDATA[Richard MacManus]]></dc:creator><pubDate>Wed, 01 Jul 2026 00:20:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FQL_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FQL_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FQL_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 424w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 848w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FQL_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg" width="1280" height="960" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:960,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:940082,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.latent.space/i/204364759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FQL_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 424w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 848w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!FQL_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf8be96e-5c79-4412-baaa-e987da5ef53a_1280x960.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Sierra&#8217;s Natalie Meurer at the AI Engineer World&#8217;s Fair today.</figcaption></figure></div><p><a href="https://www.linkedin.com/in/nataliemeurer/">Natalie Meurer</a> is Head of Agent Engineering at Sierra, where she leads a global team of more than 120 engineers building conversational AI agents for enterprise customer service. Before joining Sierra, she worked in technology policy, taught herself to code and spent five years at Palantir.</p><p>Forward deployed engineering (FDE) was one of the tracks running at today&#8217;s <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer World&#8217;s Fair</a>. As Meurer explained to Latent Space before the session she presented, FDE began as a model for placing highly technical employees close to customers. But the title now covers a wide range of roles across the AI industry &#8212; including what Sierra calls the <strong>agent engineer</strong>: an engineer who combines systems integration and agent development with an understanding of customer operations, product, and the end-user experience.</p><p>In this Q&amp;A, Meurer argues that FDE is defined more by accountability than by a particular skill set, adding that product and customer-facing engineering may be starting to converge.</p><h2>Defining forward deployed engineering</h2><p><strong>Latent Space:</strong> What is your definition of a forward deployed engineer?</p><p><strong>Natalie Meurer:</strong> That is really the point of my session: the role lacks a consistent definition.</p><p>If you look at its historical trajectory through to the present, it is more clearly defined by accountability to customers than by the shape of the role or the work you are doing.</p><p>There is power in having that accountability. But the range of associated skill sets has become so broad that it can almost become nonsensical.</p><p><strong>Latent Space:</strong> How did you get into this kind of role?</p><p><strong>Meurer:</strong> I began in technology policy. I was a policy nerd who learned to code on the side, which earned me a role as an engineer on the privacy team at Palantir.</p><p>I spent about five years there, working across law enforcement, defence and infrastructure engineering. I then went to business school because I wanted to bring the business dimension into the mix. After that, I joined Sierra and founded the agent engineering function.</p><h2>Why Sierra calls them agent engineers</h2><p><strong>Latent Space:</strong> Did Palantir&#8217;s forward deployed engineering model influence the role at Sierra?</p><p><strong>Meurer:</strong> Somewhat, although we intentionally called the role <strong>agent engineer</strong>, rather than forward deployed engineer.</p><p>Forward deployed engineering can mean so many things. We thought the title should capture the shape of the technical work, rather than only the customer-obsession element. That is why we chose agent engineer.</p><p>I see agent engineering as either a subset of, or adjacent to, forward deployed engineering. It describes a more specific form of customer-facing engineering focused on developing agents.</p><h2>What an agent engineer does</h2><p><strong>Latent Space:</strong> What does your team do when working with a customer?</p><p><strong>Meurer:</strong> Sierra builds conversational AI agents for inbound and outbound customer service. Our work includes integrating customer systems with low-latency voice and chat agents, as well as agents that operate over email.</p><p>The role requires technical skills such as data integration, but it also requires taste. You need to understand what sounds good and what will feel human when you are designing a voice agent. That element is particular to agent engineering.</p><p><strong>Latent Space:</strong> Does an engagement begin with a defined use case, or do you help the customer decide what to build?</p><p><strong>Meurer:</strong> We conduct discovery with our customers. We try to find the intersection between problems that are genuinely difficult &#8212; because we are good at difficult problems &#8212; and problems that will have a meaningful business impact.</p><p>In financial services, for example, that might begin with dispute processing. It is complex and needs to be done correctly, but it is also a high-emotional-intelligence interaction. If somebody sees a fraudulent charge on their credit card statement, they may be frightened, and the agent needs to calm them down.</p><p>Almost every Sierra customer is also somewhere on the trajectory towards using an agent as its front-door interactive voice response system: the first entity that answers when a customer calls.</p><h2>The hard work is often above the model layer</h2><p><strong>Latent Space:</strong> How much of the work involves the underlying AI models?</p><p><strong>Meurer:</strong> We think of our agents as an orchestrated constellation of models. Internally, we are constantly evaluating the best model for a particular job, and we bring the best of that work to our customers.</p><p>In practice, most customer-specific work takes place at the orchestration layer rather than in the models themselves. We sometimes integrate with a customer&#8217;s own models, and we also help customers use the platform and build agents themselves.</p><p>A lot of the work involves helping them apply their internal knowledge and context.</p><h2>Custom deployments and reusable patterns</h2><p><strong>Latent Space:</strong> How much of the work is customer-specific, and how much can be reused?</p><p><strong>Meurer:</strong> It is a mixture of both.</p><p>Every customer is building an agent that is intentionally specific to its organization. It should represent the best possible interaction with that particular brand.</p><p>Other capabilities are more reproducible. Answering questions from a knowledge base, for example, is a fairly universal problem. We also have industry experts across financial services, healthcare, travel and hospitality, and retail who bring domain knowledge and best practices.</p><p>But the fundamental appeal of what we are selling is something custom. We have seen large organizations across industries reach production in as little as 40 to 60 days.</p><p>Each agent is still customized around the customer&#8217;s APIs, systems, standard operating procedures, brand and tone.</p><h2>Agents as enterprise systems</h2><p><strong>Latent Space:</strong> Is agent development becoming primarily an orchestration problem?</p><p><strong>Meurer:</strong> There are many different flavours of multi-agent architecture. The term &#8220;agent&#8221; can refer to the entity that answers the phone, but it can also refer to a sub-agent or even a single prompt equipped with tools.</p><p>Every enterprise we work with wants to know how it can maintain everything its agentic ecosystem is capable of doing. It needs to manage all the integrations and all the teams that contribute to the agent.</p><p>Part of that is a change-management problem.</p><p>At Sierra, we tend to think of a single agent as managing the entire customer interaction, regardless of the particular subtask involved. We call those subtasks <strong>journeys</strong>.</p><p>Enterprises nevertheless need a way for hundreds or thousands of people to contribute to these systems, understand what is changing and follow a discrete release process.</p><h2>Product engineering and FDE are converging</h2><p><strong>Latent Space:</strong> As companies develop more internal expertise, how will the FDE role evolve?</p><p><strong>Meurer:</strong> I think it will remain customer-facing. But when code becomes cheap to author, it also becomes easier to translate customer insights directly into a product.</p><p>Product engineering and forward deployed engineering are therefore converging in some respects &#8212; at least among the best people in each role.</p><p>If you are a product engineer, you should be talking to customers. If you are a forward deployed engineer, you should be building the product. I think that is new.</p><p>Being customer-facing will remain important. Even if you had an AGI-like reasoning model that could work out how to perform a process each time, you would still need to encode that process appropriately.</p><p>You do not want the system independently figuring out how to handle an order return for the 100,000th time that week. You want a consistent process that it follows.</p><p>That makes customer service different from some other agentic use cases. A coding agent is often trying to solve a new problem for the first time. In customer service, you are solving essentially the same problem, framed slightly differently, perhaps 100,000 times a week.</p><p>That creates a different need for both the platform and the partner helping the customer encode its rules. Agents will become easier to build, but there will always be a place for people who can work with customers and translate what they learn into the product.</p><h2>Why generalists may become more valuable</h2><p><strong>Latent Space:</strong> Will developers increasingly need product and customer-facing skills?</p><p><strong>Meurer:</strong> That is my belief. I think the best developers will develop those skills.</p><p>Many people are asking what the engineering role will look like in one or two years. One view is that specialists will become even more important because they possess knowledge that is not readily available to an agent.</p><p>The other view, which I lean towards, is that generalists will become more valuable.</p><p>Forward deployed engineering has historically been the classic generalist role because it combines engineering with the customer-facing nature of the job.</p><p>Forward deployed engineers &#8212; or agent engineers &#8212; therefore inhabit one of the most forward-looking areas in AI and engineering.</p><p><strong>Latent Space:</strong> Could &#8220;agent engineer&#8221; eventually become the default term?</p><p><strong>Meurer:</strong> I am not sure. I expect engineering as a whole to move towards a more holistic definition, one that may incorporate more of what we currently call forward deployed engineering.</p><p>The market currently has go-to-market engineers, forward deployed engineers, agent engineers and AI engineers.</p><p>I think all of those will become different parts of the engineering craft. We will also discover entirely new jobs for engineers to do.</p>]]></content:encoded></item></channel></rss>