The Best Model Is a Routing Decision
The best model isn't the smartest — it's whichever should get the next unit of work. Notes from NVIDIA on routing, local models, and the AI factory.
Co-founder and CTO of Speedscale, expert in Agentic AI and cloud data warehousing. • 9 posts published
The best model isn't the smartest — it's whichever should get the next unit of work. Notes from NVIDIA on routing, local models, and the AI factory.
Our v2 release looked clean in HTTP tests until we diffed the SQL workload — an N+1 loop, a startup migration, and 70 ms of extra DB time hiding in plain sight. Here's how to compare two releases without database access.
The trace was sampled out. I found the bug anyway — by filtering recorded traffic on the customer's email instead of a trace ID. Here's how to follow one request across four services with no trace IDs and no OpenTelemetry.
Metrics, logs, and traces were built for humans and cheap storage. AI inverts both assumptions, and the next maturity level is a deterministic replay sandbox.
Logs, metrics, and traces are a lossy compression of production. Five things you can do with a traffic data lake that observability can't.
A Kubeshark alternative that goes beyond observability. Stream live cluster traffic into proxymock, then replay or mock it locally from your laptop.
Export recorded proxymock traffic to Datadog Synthetics in one command. Auth headers redacted, global variables created. No scripting, no flaky journeys.
UI synthetics only tell you something is broken. Traffic replay per microservice isolates failures before any human walks up. Zero scripts required.
SaaS AI fails when agents need continuous access to your codebase and internal APIs. Here's why BYOC is the only deployment model that works at scale.
Choose the desktop proxymock or the hosted cloud trial to get started.