395 episodios
- As AI coding agents accelerate software development, they also create new challenges for site reliability engineers (SREs), who are increasingly responsible for debugging systems that no single human fully understands. In this episode ofThe New Stackpodcast, Sam Farid and Nate Heinrich of Chronosphere argue that AI agents should also be used for root-cause analysis, helping teams diagnose failures more quickly as model capabilities continue to improve.
Rather than immediately purchasing a commercial solution, they recommend organizations first build an in-house AI SRE. The process of documenting systems, dependencies, and operational knowledge creates valuable context that enables AI agents to troubleshoot effectively while improving institutional knowledge. Although Chronosphere offers its own AI SRE platform, the hosts emphasize that building an internal prototype helps teams understand their needs before evaluating vendor tools. As AI-generated code becomes more common, organizations that invest in mapping their systems and leveraging AI for operations will be better equipped to reduce downtime and support increasingly complex software environments.
Learn more from The New Stack around AI SREs:
5 ways SRE AI agents are set to augment human capabilities
The Future of AI in SRE: Preventing Failures, Not Fixing Them
AI Reliability Engineering: Welcome to the Third Age of SRE
Join our community of newsletter subscribers to stay on top of the news and at the top of your game. - In this episode with The New Stack Agents, Frederic Lardinois, NVIDIA’s Joey Conway says advances in AI over the past year have dramatically improved the capabilities of local models, making them practical for enterprise and personal use alongside frontier cloud models. Rather than replacing large models, Conway envisions a “system of models” where specialized local models handle routine, cost-sensitive, or privacy-focused tasks, while larger frontier models tackle more complex reasoning. He explains that organizations can fine-tune smaller open models using domain-specific data, creating expert AI agents that reflect the specialized roles found within businesses.
NVIDIA supports this ecosystem through open models, training tools, and software such as NeMo, Dynamo, and Nemotron. Conway also highlights the growing importance of agentic harnesses, which give AI models access to tools, memory, and iterative workflows, significantly improving performance and reducing costs. Looking ahead, he expects AI orchestration to become increasingly important, with intelligent routing systems selecting the right model for each task based on complexity, cost, latency, and data governance requirements, enabling enterprises to balance performance, security, and efficiency.
Learn more from The New Stack around NVIDIA's latest updates in AI:
Palantir and Nvidia want to change who owns government AI
Nvidia's best model is now live
Join our community of newsletter subscribers to stay on top of the news and at the top of your game. - In this episode, Mark Russinovich, CTO of Microsoft Azure revealed Brain, the AI-powered AIOps system that continuously monitors Azure’s health, detects incidents, identifies root causes, and increasingly automates responses such as pausing problematic deployments and notifying affected customers. Built on Azure Resource Graph, Brain creates a real-time digital twin of Azure, mapping dependencies across hundreds of services, data centers, and regions. Although Brain predates the generative AI boom, years of data engineering, standardized service-level indicators (SLIs), and machine learning laid the foundation for today’s capabilities.
Brain combines standardized SLIs, service-specific monitoring, and third-party signals to detect anomalies, while ML models dynamically establish service baselines and correlate outages with software rollouts. Microsoft says automated notifications have reduced customer support tickets by four to six times, with 80–90% of Brain-covered services receiving notifications within 15 minutes, often in under five. The company is also layering LLM-powered agents, called Triangle, on top of Brain to streamline incident routing and eventually enable AI agents to autonomously troubleshoot and remediate outages.
Learn more from The New Stack around the latest in Microsoft Azure:
Meet Brain, the AI that decides when Azure is officially down
Microsoft's pitch to enterprises: Ditch Azure Repos for GitHub, despite its rocky reliability record
Join our community of newsletter subscribers to stay on top of the news and at the top of your game. - Subquadratic is beginning to back up its ambitious claims with benchmarks and third-party validation for its SubQ 1.1 Small model, which uses its proprietary Sparse Attention (SSA) architecture to dramatically improve long-context performance. Rather than comparing every token to every other token, SSA selectively processes relationships, enabling near-linear scaling while maintaining high accuracy across context windows of up to 12 million tokens. The company reports near-perfect retrieval performance, competitive coding and reasoning benchmarks, and compute savings of up to 1,000x at maximum context lengths.
Rather than targeting frontier models immediately, Subquadratic is focusing on enterprise customers that need efficient analysis of massive datasets. The current model was built by replacing the dense attention mechanism in an existing open-weight model and then continuing long-context pretraining. Looking ahead, the startup plans to release a larger mid-tier model while continuing research into "zero attention" architectures that could eliminate attention mechanisms altogether, with the long-term goal of surpassing today's transformer-based AI models in both efficiency and capability.
Learn more from The New Stack around cloud spending:
The context window has been shattered: Subquadratic debuts a 12-million-token window
What comes after attention? This startup says it already knows.
Join our community of newsletter subscribers to stay on top of the news and at the top of your game. “The harness is where the hard work is”: Harness bets on agents that enterprises can trust in production
02/07/2026 | 19 minHarness has introduced Autonomous Worker Agents, a new capability that allows enterprises to replace rigid CI/CD pipeline scripts with AI agents that can deploy applications, run tests, and perform security scans while operating under existing governance, security, and audit controls. Unlike Harness' existing expert agents, which assist developers with coding and pipeline creation, Worker Agents autonomously execute pipeline tasks within customer-controlled infrastructure. Agents are defined using simple Markdown files, draw context from the Harness Software Delivery Knowledge Graph, and run in sandboxed environments with scoped permissions and policy enforcement.
Harness also provides built-in audit trails that record prompts, decisions, and outcomes, along with token budgets and approval gates to control AI costs. The launch includes an Agent Marketplace featuring Harness-managed, certified partner, and community-built agents. CEO Jyoti Bansal said production AI agents require far stronger safeguards than coding assistants, positioning Harness' governance and knowledge graph as key differentiators. Looking ahead, the company envisions fully autonomous software engineering, where AI agents manage the software lifecycle while humans oversee high-risk decisions.
Learn more from The New Stack around AI software delivery:
AI won't speed up software delivery - nothing has
How to solve the AI paradox in software development with intelligent orchestration
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Más podcasts de Noticias
Podcasts a la moda de Noticias
Acerca de The New Stack Podcast
The New Stack Podcast is all about the developers, software engineers and operations people who build at-scale architectures that change the way we develop and deploy software.
For more content from The New Stack, subscribe on YouTube at: https://www.youtube.com/c/TheNewStack
Sitio web del podcastEscucha The New Stack Podcast, 6AM W y muchos más podcasts de todo el mundo con la aplicación de radio.net

Descarga la app gratuita: radio.net
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app
Descarga la app gratuita: radio.net
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app


The New Stack Podcast
Escanea el código,
Descarga la app,
Escucha.
Descarga la app,
Escucha.
The New Stack Podcast: Podcasts del grupo





























