archive
How we run thousands of isolated agent trials with smolvm
Engineering
The Modalities of Testing for AI Agents
Insights
Replaying Production Conversations Safely
Tutorials
Your AI Agent's First 10,000 Failures Should Be Free
Perspective
Why Eval Suites Go Stale
Opinion
Scenario Discovery from Live Traffic
AI
Selling Agent Reliability to the Enterprise
B2B
From 10 to 10,000 Simulated Sessions
Scaling
Grading Rubrics that Survive Model Swaps
Tutorials
The Recovery Loop: Fail, Learn, Redeploy
Insights
Agents Do Not Fail Randomly
AI
Pricing Reliability as a Feature
B2B
Load Patterns that Break Tool-Calling
Scaling
Stop Shipping Agents on Vibes
Opinion
AI fails when conditions change. Make sure yours is ready.
Connect your tools, capture real conversations, and start replaying events for training and evaluation - all from your dashboard.
