The 0% error claim at 1B tokens/day is the number that deserves unpacking. At that scale, "0%" in practice usually means below the detection threshold or the error definition is narrowly scoped. What's the observable unit here — failed API calls, degraded responses, or something else? Genuinely curious how they defined and measured it, because getting that right at this scale is the actual hard problem.
The disposable software framing clicked for me. I spent two months building a custom task board for my AI agent. Last week replaced the whole thing with an open-source tool (Fizzy) in a few hours. The custom code wasn't bad - it just didn't need to exist. The spec-as-documentation approach is basically what I've landed on too: agent reads a markdown file, executes against it, updates state. Not 1M LOC, not zero human code, but the ratio keeps shifting. The question I'm sitting with is where the human judgment layer should live once the agent starts modifying the spec itself.
Human health has been profoundly transformed over the past century, largely due to scientific progress and international collaboration. The global maternal mortality rate has fallen by more than 40% since 2000, and deaths among children under five have been reduced by over 50%. Advances in technology, scientific knowledge and skills, and collaboration between different disciplines, sectors and countries continue to turn once-life-threatening health challenges – such as elevated blood pressure, cancer diagnoses or HIV infection – into manageable health issues, extending and improving lives worldwide.
“Yet, health threats continue to grow, fuelled by climate impacts, environmental degradation, geopolitical tensions and shifting demographics. These challenges include persistent diseases and strained health systems as well as emerging diseases with epidemic or pandemic potential. Across the globe, thousands of scientists – together with organizations such as WHO – are accelerating research and developing policies, tools and innovations needed to protect communities today and safeguard the health of future generations.
“Science is one of humanity’s most powerful tools for protecting and improving health,” “People in every country live longer and healthier lives on average today than their ancestors did, thanks to the power of science. Vaccines, penicillin, germ theory, MRI machines and the mapping of the human genome are just some of the achievements that science has delivered that have saved lives and transformed health for billions of people.”
WHO has been coming through and saving life for 78yrs and still counting..
If you look at his tweet reproduced in the article, the guy says "post-merge code review" - not "0% human review."
And since we KNOW that 45% of AI-generated code is INSECURE code, this tells us that OpenAI clearly does not give a shit about the quality of their code.
This is "agile programming" taken to its logical conclusion: we don't give a shit about the code as long as we can satisfy management by churning out massive amounts of lines of code.
That used to be the sign of a dysfunctional development organization - before management decided that lines of code was the only management metric needed.
This is why Windows has umpteen million lines of code - and is next to totally insecure.
This is why I despise Altman and his crew of lunatics. They are bullshit artists.
The 0% error claim at 1B tokens/day is the number that deserves unpacking. At that scale, "0%" in practice usually means below the detection threshold or the error definition is narrowly scoped. What's the observable unit here — failed API calls, degraded responses, or something else? Genuinely curious how they defined and measured it, because getting that right at this scale is the actual hard problem.
The disposable software framing clicked for me. I spent two months building a custom task board for my AI agent. Last week replaced the whole thing with an open-source tool (Fizzy) in a few hours. The custom code wasn't bad - it just didn't need to exist. The spec-as-documentation approach is basically what I've landed on too: agent reads a markdown file, executes against it, updates state. Not 1M LOC, not zero human code, but the ratio keeps shifting. The question I'm sitting with is where the human judgment layer should live once the agent starts modifying the spec itself.
Human health has been profoundly transformed over the past century, largely due to scientific progress and international collaboration. The global maternal mortality rate has fallen by more than 40% since 2000, and deaths among children under five have been reduced by over 50%. Advances in technology, scientific knowledge and skills, and collaboration between different disciplines, sectors and countries continue to turn once-life-threatening health challenges – such as elevated blood pressure, cancer diagnoses or HIV infection – into manageable health issues, extending and improving lives worldwide.
“Yet, health threats continue to grow, fuelled by climate impacts, environmental degradation, geopolitical tensions and shifting demographics. These challenges include persistent diseases and strained health systems as well as emerging diseases with epidemic or pandemic potential. Across the globe, thousands of scientists – together with organizations such as WHO – are accelerating research and developing policies, tools and innovations needed to protect communities today and safeguard the health of future generations.
“Science is one of humanity’s most powerful tools for protecting and improving health,” “People in every country live longer and healthier lives on average today than their ancestors did, thanks to the power of science. Vaccines, penicillin, germ theory, MRI machines and the mapping of the human genome are just some of the achievements that science has delivered that have saved lives and transformed health for billions of people.”
WHO has been coming through and saving life for 78yrs and still counting..
If you look at his tweet reproduced in the article, the guy says "post-merge code review" - not "0% human review."
And since we KNOW that 45% of AI-generated code is INSECURE code, this tells us that OpenAI clearly does not give a shit about the quality of their code.
This is "agile programming" taken to its logical conclusion: we don't give a shit about the code as long as we can satisfy management by churning out massive amounts of lines of code.
That used to be the sign of a dysfunctional development organization - before management decided that lines of code was the only management metric needed.
This is why Windows has umpteen million lines of code - and is next to totally insecure.
This is why I despise Altman and his crew of lunatics. They are bullshit artists.