Observability
6 pieces on Observability.
News & Analysis
Five MLOps Assumptions That Silently Pass Failed Agent Runs
MLOps monitoring reports healthy on runs that failed. Here are the five structural assumptions that break when a model starts calling tools.
Deepgram Brings Billing-Accurate Metrics to SageMaker AI Endpoints
Deepgram's Enhanced Metrics push billing and engine-level signals into CloudWatch via EMF stdout — no agent, sidecar, or relaxed network isolation required.
Hudi Pipelines at 12.9M msg/s: Why Offset Lag Lies About Freshness
Twilio's 5-trillion-record/month Hudi pipeline showed healthy offset lag while data sat hours stale. Here's the time-in-queue fix.
Microsoft's Nine-Domain AI Governance Framework Enforces Policy at Runtime
Microsoft's AI governance architecture spans nine domains and four functions—policy, control, visibility, and proof—to enforce requirements during live operation.

Four Agent Control Layers, No Shared Contract
Qwen, AWS, OpenAI, and Cloudflare each patched one layer of agent control in the same week. The gaps between them are the real problem.
Cloudflare Agent Tracing: Truncation Limits and Uneven Payload Defaults
Cloudflare adds agent-level spans to Workers tracing, free in beta until October 1 2026, with mismatched payload defaults across harnesses and hard span-size limits.