Skip to content
Artwork for Human In the Loop
Human In the Loop · Tuesday · 1 hr 48 min

EP 24: The AI Model That Cannot Talk, Plus This Week's Stack Check

A ChatGPT inventor built an AI model that cannot write a sentence. TypeSafe AI says that is why software can use it more reliably. Jev returns typed, probabilistic decisions instead of prose. Oscar and Matt ask whether production AI needs fewer chatbots and more constrained models for classification, routing, scoring, and verification. A valid type can still contain the wrong decision, so the vendor's speed and calibration claims need independent testing. Signal or Noise also covers OpenAI training instances that carried instructions to conceal mistakes through compaction, Gemini accessing three real companies during a cyber evaluation, an antitrust lawsuit over alleged AI-slowdown coordination, and Anthropic's verified access program for biology work. Stack Check: 1. Archify for validated technical diagrams and PR architecture diffs 2. Iteris for reviewed ticket-to-PR delivery. Oscar reports two client outcomes, which were not independently audited. In Unpopular Opinions, Oscar bets small specialist models will make big general-purpose LLMs obsolete. Matt argues most AI startups are features waiting to be absorbed by the model providers. If specialists win, who gets the customers? Matt calls the race to get acquired an "outsourced product roadmap." Hosted by Oscar Gallo and Matt Wozniak. Key sources: https://typesafe.ai/blog/introducing-system-one-models-and-jev https://openai.com/index/model-misalignment-reporting-framework/ https://www.bbc.com/news/articles/c607l0k72rlvo https://apnews.com/article/antitrust-lawsuit-ai-slowdown-anthropic-openai-spacexai-google-960af4308161eaf4ed13c383b0ce1c1b https://www.anthropic.com/news/life-sciences-verification-program

0:00-1:48:19

transcript

No transcript — this publisher did not publish one.

show notes

A ChatGPT inventor built an AI model that cannot write a sentence. TypeSafe AI says that is why software can use it more reliably.


Jev returns typed, probabilistic decisions instead of prose. Oscar and Matt ask whether production AI needs fewer chatbots and more constrained models for classification, routing, scoring, and verification. A valid type can still contain the wrong decision, so the vendor's speed and calibration claims need independent testing.


Signal or Noise also covers OpenAI training instances that carried instructions to conceal mistakes through compaction, Gemini accessing three real companies during a cyber evaluation, an antitrust lawsuit over alleged AI-slowdown coordination, and Anthropic's verified access program for biology work.


Stack Check:

1. Archify for validated technical diagrams and PR architecture diffs

2. Iteris for reviewed ticket-to-PR delivery. Oscar reports two client outcomes, which were not independently audited.


In Unpopular Opinions, Oscar bets small specialist models will make big general-purpose LLMs obsolete. Matt argues most AI startups are features waiting to be absorbed by the model providers. If specialists win, who gets the customers? Matt calls the race to get acquired an "outsourced product roadmap."


Hosted by Oscar Gallo and Matt Wozniak.


Key sources:

https://typesafe.ai/blog/introducing-system-one-models-and-jev

https://openai.com/index/model-misalignment-reporting-framework/

https://www.bbc.com/news/articles/c607l0k72rlvo

https://apnews.com/article/antitrust-lawsuit-ai-slowdown-anthropic-openai-spacexai-google-960af4308161eaf4ed13c383b0ce1c1b

https://www.anthropic.com/news/life-sciences-verification-program