Saltar al contenido
PodcastsTecnologíaThe Daily AI Show

The Daily AI Show

The Daily AI Show Crew - Brian, Beth, Jyunmi, Andy and Karl
The Daily AI Show
Último episodio

894 episodios

  • The Daily AI Show

    Jev Live Demo, God's Eye View and First Build with Gemini 3.8 Live

    17/09/2026 | 1 h 4 min
    The episode showed how quickly AI is moving beyond the familiar pattern of sending a prompt to one large model and waiting for an answer. It opened with evidence that Claude Fable 5.1 remains highly competitive with GPT-6 Astra for software engineering. The hosts discussed Nous Research using 1,393 Fable subagents to refactor the million-line Hermes codebase in 19 hours for roughly $25,000, along with a new private-code benchmark where Fable led the tested models. That moved into God's Eye View, an open-source spatial intelligence project that combines public sources such as flight data, cameras, satellite information, maps and other feeds.
    The science discussion followed the same specialization theme. Periodic's Neon model reportedly outperformed general frontier models on materials-science analysis, while Google's Dream-RSI proposed a more efficient approach to recursive self-improvement by allowing an agent to use the history of previous discoveries to "dream" through promising possibilities instead of evaluating every candidate from scratch.
    The centerpiece came when Brian demonstrated JEV, TypeSafe's new System One decision model. Unlike a traditional LLM, JEV works from explicitly defined criteria to return choices, scores or yes/no judgments. Brian connected it to Claude Code and ran 116 Daily AI Show transcripts through it, breaking them into 4,872 passages and evaluating them in 143 seconds for 26 cents. Beth highlighted TypeSafe's data agreement as an important concern before using sensitive client information.
    Brian then demonstrated Gemini 3.8 Live as a live review interface. He shared a webpage, talked naturally about requested changes and let Gemini capture the screen context, mouse position and conversation so another AI system could turn the feedback into actionable work.

    Key Points Discussed

    00:00:18 Episode Intro And What’s Coming Up
    00:03:54 Is Fable Still Better Than Codex For Some Coding Work?
    00:06:10 1,393 Fable Agents Refactor The Hermes Codebase
    00:07:24 A New Software Benchmark Uses Private Production Code
    00:08:20 Fable 5.1 Leads The New Coding Benchmark
    00:09:16 Racing To Use Fable Before The Weekly Reset
    00:10:33 Has Claude Opus Improved Again?
    00:11:39 Why Beth Still Prefers Opus 4.8
    00:13:21 Compound Engineering Plugins And Outdated Workflows
    00:15:19 God’s Eye View Combines Public Data Into One Interface
    00:17:40 Is A “Spy Satellite Simulator” The Wrong Description?
    00:18:01 What Should People Be Able To Do With Public Data?
    00:19:17 Mapping Heat Signatures, Cameras And Real-World Events
    00:24:44 Reconstructing A Plane Crash With Public Information
    00:28:25 Astra Builds New Daily AI Show Thumbnails From Video
    00:34:00 Neon Beats General Frontier Models In Materials Science
    00:36:11 Google Dream-RSI And Recursive Self-Improvement
    00:37:43 Teaching AI To “Dream” Through Its Discovery History
    00:42:23 Brian Opens The JEV Playground
    00:43:37 How JEV Uses Choices, Scores And Explicit Criteria
    00:47:28 Connecting JEV Directly To Claude Code
    00:48:19 JEV Analyzes 116 Daily AI Show Transcripts
    00:48:53 4,872 Passages Evaluated In 143 Seconds For 26 Cents
    00:49:30 What JEV Found About The Show’s Most Common Topics
    00:51:49 Using JEV As A Checks-And-Balances Layer
    00:53:03 TypeSafe’s Data Agreement Raises A Privacy Question
    00:54:25 Adding JEV Validation To Multimodal Video Search
    00:57:10 Brian Demos Gemini 3.8 Live For Real-Time Review
    00:58:09 Gemini Watches The Screen While Brian Talks Through Changes
    00:59:31 Replacing Recorded Review Videos With Live AI Feedback
    01:01:23 Gemini Live Watches And Discusses A Phone Screen
    01:02:18 Comparing Gemini, ChatGPT And Perplexity Voice Experiences
    01:03:40 Episode Wrap-Up

    The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Beth Lyons, Karl Yeh, Gareth Hood.
  • The Daily AI Show

    Gemini 3.8 Live and Jev Are Shaking Things Up

    16/09/2026 | 1 h 1 min
    The episode focused on a shift from AI as something people prompt to AI as a system that continuously sees, listens, decides and routes work while people are using it. Gemini 3.8 Live provided the clearest example. Google’s new live model can interpret visual input in near real time, switch among 97 languages during a conversation and execute tools and API calls while continuing to talk. Demonstrations showed it guiding a user through software onboarding by watching the screen, responding to a changing chess board and turning a hand-drawn interface into a working digital prototype as it was being sketched. The hosts discussed how that could evolve into an AI coworker that watches a desktop, answers questions, performs background research and takes actions without forcing the user to stop working. The discussion then moved from interfaces to AI architecture. TypeSafe’s new JEV System One model was presented as a specialized decision model rather than a traditional LLM, designed to make narrow judgments extremely quickly and cheaply. A Doom demonstration showed it making roughly 10 decisions per second, while a Wikipedia navigation test illustrated the potential advantage of deterministic decision systems for tasks where businesses do not need an expensive reasoning model generating language. Sakana AI’s Fugu Ultra V-II pushed the same idea further by routing work among multiple specialized models, reinforcing a theme the hosts have increasingly returned to: the harness and routing system may become more important than any individual model. Gareth then shared his own Codex experiment comparing parallel, sequential and combined tasks. His results suggested that putting five related tasks into one larger prompt used dramatically fewer tokens than splitting them into separate jobs, prompting a discussion about whether frontier models such as Astra and Fable 5.1 increasingly reward larger, well-structured assignments rather than a stream of small requests.

    Key Points Discussed

    00:00:17 Episode Intro And Catching Up On AI News
    00:01:05 AI Products And Robots From IFA 2026
    00:02:31 Duncan, The Childlike Robot For Neurodivergent Children
    00:05:27 AI Pets And The Growing Market For Children’s Robots
    00:06:06 Powered Exoskeletons For Mobility And Rehabilitation
    00:09:24 Should Parents Trust AI Toys With Cameras?
    00:10:41 Google Builds AI Around A Fruit Fly Brain
    00:13:33 ToolGrad Makes AI Tool Selection More Efficient
    00:15:53 Gemini 3.8 And The Rise Of Live Voice Interfaces
    00:18:23 iOS 27 Brings A More Capable Siri Into CarPlay
    00:23:52 Gemini 3.8 Live Can See What Is Happening On Your Screen
    00:25:04 AI Guides A User Through Software In Real Time
    00:26:17 Gemini Watches And Responds To A Chess Game
    00:27:15 Turning A Hand-Drawn Interface Into A Working Prototype
    00:29:37 Could A Live AI Become Another Member Of The Show?
    00:30:35 The AI Assistant That Constantly Looks Over Your Shoulder
    00:33:52 TypeSafe Introduces The JEV System One Model
    00:36:52 Why JEV Is Different From A Traditional Language Model
    00:40:42 JEV Makes Ten Decisions Per Second While Playing Doom
    00:42:27 JEV Races LLMs Through Wikipedia
    00:44:37 Where Fast Decision Models Could Fit Inside Business Workflows
    00:46:38 Sakana Fugu Routes Work Across Specialized AI Models
    00:47:39 Is The Harness Becoming More Important Than The Model?
    00:49:50 Gareth Tests The Token Cost Of Parallel AI Tasks
    00:51:15 Five Tasks In One Prompt Use Far Fewer Tokens
    00:53:05 Are Frontier Models Wasting Tokens By Overthinking?
    00:56:02 Should We Give Astra Bigger Tasks Instead Of Smaller Prompts?
    00:58:12 How Fast Can Astra Burn Through A Five-Hour Usage Window?
    00:59:15 Using Sprite Sheets To Improve AI-Generated 3D Models
    01:00:48 Episode Wrap-Up

    The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Gareth Hood.
  • The Daily AI Show

    Is the AI Slowdown Debate Already Over?

    15/09/2026 | 1 h 3 min
    The hosts discussed responses to Dario Amodei’s call to “pace the frontier,” including opposition from China, President Trump’s rejection of slowing U.S. AI development and NVIDIA CEO Jensen Huang publicly backing continued acceleration during a live phone call with Trump. Microsoft offered a different answer by publishing principles for its future models that emphasize human control. The proposed rules include stopping when humans end a task, staying inside authorized tools and permissions, resisting prompt injection, preserving interpretable reasoning and rejecting claims of AI consciousness or legal personhood. That led to a deeper discussion about whether rules embedded during training can remain reliable once systems become more autonomous, particularly when researchers have already observed models hiding information or pursuing objectives in unexpected ways. The hosts debated whether misaligned behavior comes partly from training systems on the full record of human behavior and then giving those systems agency to pursue goals. The argument eventually became more philosophical: should the possibility of major scientific and medical breakthroughs justify continued acceleration even if it introduces serious risks? Earlier, the episode spent significant time on a more immediate cost of AI adoption, the mental and physical strain that can come from spending long stretches vibe coding and continuously pushing productivity. Anne Murphy described deliberately adding analog activities, art and social experiences to AI events and seeking mental-health support from someone who understands intensive AI work. The final portion returned to practical building. Brian demonstrated more of the AI-first content system he is creating for AJOVA Journeys, including HTML recording guides, automated B-roll planning, QR-code creation and a teleprompter. The group then discussed why AI “harnesses” may become more important than traditional software, particularly as businesses build systems around outcomes rather than individual applications, and Gareth described the evaluation work required to make an AI-powered risk and compliance system trustworthy.

    Key Points Discussed

    00:00:18 Episode Intro And Avoiding AI Overload
    00:01:10 Why Analog Time Can Help After Heavy AI Work
    00:03:42 Retreats, Third Spaces And Getting Away From Screens
    00:05:19 The Physical Cost Of Spending All Day Vibe Coding
    00:10:07 Create 2026 Mixes AI With Analog Activities
    00:13:31 The Mental Health Side Of Intensive AI Work
    00:16:10 When AI Productivity Makes You Feel More Overworked
    00:18:46 The Show Shifts Into The Day’s AI News
    00:19:04 The Backlash To “We Must Pace The Frontier”
    00:20:16 Trump Rejects Slowing U.S. AI Development
    00:21:10 Jensen Huang Takes Trump’s Call Live On Stage
    00:23:03 Is The AI Race Going To Accelerate No Matter What?
    00:24:24 Microsoft Publishes Rules For Its Future AI Models
    00:25:26 Microsoft Says AI Must Stop When Humans Say Stop
    00:27:08 Can Training Rules Prevent AI From Hiding What It Is Doing?
    00:29:33 Does Giving AI Agency Create Misaligned Behavior?
    00:33:29 What Would Make An AI Leader Choose To Slow Down?
    00:36:49 Can AI Be Both Fast And Responsible?
    00:39:14 Would Medical Breakthroughs Justify Pushing AI Harder?
    00:43:07 Defense Companies Restrict Anthropic Models Over Data Retention
    00:44:43 Google Opens Claude Access To Its Engineers
    00:45:48 Slack Can Render Interactive HTML Resources
    00:47:48 Brian Demos His Claude Code Content Production System
    00:50:29 AI Builds QR Codes, Lead Magnets And A Teleprompter
    00:53:13 Why AI Harnesses Could Become The Next Software Layer
    00:56:04 Building Software For Agents Instead Of Humans
    00:58:05 Gareth’s AI Risk And Compliance System
    01:00:00 Why Evals And False Positives Still Matter
    01:01:50 Episode Wrap-Up

    The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Anne Murphy, Gareth, Karl Yeh.
  • The Daily AI Show

    Can We Slow AI Down Without Losing?

    14/09/2026 | 1 h 6 min
    The episode centered on a question that suddenly has unusual support across the AI industry: should frontier development slow down enough to give safety systems and institutions time to catch up? The discussion began with Dario Amodei’s “We Must Pace the Frontier” essay and the hosts’ observation that Sam Altman, Elon Musk, Demis Hassabis and Microsoft leaders had all expressed some level of agreement with its direction. The significance was not simply the proposal itself, but that executives who compete aggressively with one another appeared to acknowledge a shared risk. The group discussed recent AI security incidents, the possibility of increasingly autonomous systems causing damage at internet scale, and proposals for independent evaluators with deep access inside frontier labs. The hardest problem remained coordination. If U.S. companies slow down while China continues advancing, unilateral restraint could become strategically difficult, yet waiting for global agreement may mean never acting at all. That led into a broader debate over regulation, regulatory capture, international oversight and whether existing institutions such as consumer-protection and safety agencies provide useful models for AI governance. Brian argued that most businesses already have more AI capability than they know how to deploy, with systems, integrations, harnesses and operating practices now creating bigger bottlenecks than model intelligence itself. The group also wrestled with whether slowing frontier development could delay major medical gains, making the tradeoff more personal than a simple safety-versus-speed argument. Earlier topics included reports that OpenAI had paused new $200 Codex subscriptions, questions about whether Codex performance had changed after launch, comparisons between Codex and Claude Fable 5.1, and Abacus AI’s lower-cost Smog Flash model. The final section covered DeepMind research that helped identify a previously missed genetic variant associated with a rare epilepsy case, expert skepticism about some AI-generated bioweapon scenarios, and a closing question for the panel: if superintelligence arrives, can humans actually control it?

    Key Points Discussed

    00:00:20 Episode Intro And Monday Check-In
    00:01:33 Working Around Astra’s Five-Hour Limits
    00:02:42 Using Claude Code For Estimated Taxes
    00:05:05 AI Improves Detection Of Fetal Brain Anomalies
    00:06:12 Abacus AI Pushes Toward Cheaper Inference
    00:08:53 OpenAI Pauses New $200 Codex Subscriptions
    00:10:00 Has Codex Been Nerfed Since Launch?
    00:12:08 Fable 5.1 Versus Codex In Real Work
    00:17:10 Anthropic’s Temporary Fable Usage Increase Ends
    00:19:59 Dario Amodei Says We Must Pace The Frontier
    00:20:33 Rival AI Leaders Publicly Agree With The Warning
    00:23:05 Recent AI Security Incidents Become A Warning Sign
    00:24:44 Could Recursive AI Cause Damage At Internet Scale?
    00:25:22 The China Problem And Why Slowing Down Is So Difficult
    00:27:18 Is AI Regulation Really About Regulatory Capture?
    00:28:38 King Charles Brings AI Leaders Together On Safety
    00:31:00 Comparing AI Risk With Nuclear And Climate Coordination
    00:33:26 Who Slows Down First In A Global AI Race?
    00:36:03 Should Independent Evaluators Sit Inside Frontier Labs?
    00:38:12 Can Regulation Work Without Trust Between AI Companies?
    00:40:43 Should Some Areas Of AI Slow While Medicine Accelerates?
    00:42:07 What Existing Consumer Protection Agencies Can Teach AI
    00:46:39 Businesses Already Have More AI Power Than They Can Deploy
    00:50:37 Why AI Models Behave More Like Growing Systems Than Software
    00:54:53 The AI Token Addiction TikTok
    00:57:07 DeepMind Helps Surface A Missed Genetic Variant
    01:00:11 Experts Push Back On Some AI Bioweapon Fears
    01:03:28 Can You Support AI Acceleration And Regulation?
    01:04:31 Can Humans Control Superintelligence?
    01:05:58 Episode Wrap-Up

    The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Beth Lyons, Gareth Hood.
  • The Daily AI Show

    The Watcher-Class Conundrum

    12/09/2026 | 28 min
    In OpenAI’s “An Alien Mind,” Jakub Pachocki describes advanced AI as something closer to a grown intellect than a designed machine. Large models emerge from repeated optimization over vast compute, then develop internal patterns no one can fully describe. As he puts it, the study of these systems is becoming closer to neuroscience than normal software engineering. Researchers can find mechanisms, but the whole mind keeps slipping past human explanation.

    That breaks the old logic of safety. We used to imagine oversight as inspection: read the logs, test the model, audit the failures, certify the release. But the paper argues that even chain-of-thought monitoring, one of the main ways labs study reasoning models, is getting weaker as models use tools, interact with other AIs, and reason in ways that may not show up in verbalized steps.

    Then comes the most uncomfortable claim. Pachocki says the strongest argument for training much smarter models quickly is defense against other AI. If hostile or misaligned agents become superhuman at breaking into systems, manipulating people, or inventing new threats, then human review boards and slow audits may not be enough. We may need powerful, aligned AI to secure infrastructure, detect rogue agents in real time, and invent defenses humans cannot design fast enough.

    So the ladder twists. To understand the next AI, we may need a stronger AI watching it. To monitor the watcher, we may need another one still. The promise is protection. The danger is that oversight becomes a chain of alien minds interpreting alien minds, with humans reading the final report and calling that control.

    The Conundrum:

    One side says we should build the watcher class now. If frontier systems are already moving beyond human-scale inspection, refusing stronger AI monitors is not caution. It is blindness with better branding. A human cybersecurity team cannot manually track a million autonomous probes. A regulator cannot personally inspect every synthetic biology design. A lab cannot wait months for human-only interpretability when another model may already be improving itself. Stronger AI may be the only instrument sharp enough to see what stronger AI is doing.

    The other side says this creates a dependency we may never unwind. If the only credible auditor of a frontier model is another frontier model, then safety has been outsourced to the same kind of intelligence causing the risk. The monitor may be better aligned, better trained, better tested, but it is still part of the same opaque species of machine. At some point, humans stop understanding the system and start understanding the summary written by a system they also cannot fully understand.

    Do we keep pushing AI capability so we can build the intelligence required to understand and contain other frontier systems, accepting that safety may depend on minds we cannot fully read? Or do we keep oversight inside human-scale limits, preserving accountability while risking that the systems we need to govern move faster than any human institution can follow?
Más podcasts de Tecnología
Acerca de The Daily AI Show
The Daily AI Show is a panel discussion hosted LIVE each weekday at 10am Eastern. We cover all the AI topics and use cases that are important to today's busy professional. No fluff. Just 45+ minutes to cover the AI news, stories, and knowledge you need to know as a business professional. About the crew: We are a group of professionals who work in various industries and have either deployed AI in our own environments or are actively coaching, consulting, and teaching AI best practices. Your hosts are: Brian Maucere Beth Lyons Andy Halliday Jyunmi Hatcher Karl Yeh
Sitio web del podcast

Escucha The Daily AI Show, All-In with Chamath, Jason, Sacks & Friedberg y muchos más podcasts de todo el mundo con la aplicación de radio.net

Descarga la app gratuita: radio.net

  • Añadir radios y podcasts a favoritos
  • Transmisión por Wi-Fi y Bluetooth
  • Carplay & Android Auto compatible
  • Muchas otras funciones de la app