Data Quality Work That Actually Sticks: separating incident response from root-cause fixes
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
39 articles
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I focused on making delivery decisions auditable and repeatable—documenting intent, success criteria, and rollback paths to reduce tribal knowledge.
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
Latency: P50, P95, P99 response times Cost: Per query, per user, per day Quality: Response relevance, accuracy
Observability is essential for AI applications. Without it, you're flying blind on costs, performance, and quality. Implement these patterns early and…
LLM systems have unique characteristics: non-deterministic outputs, variable latency, complex cost models, and quality metrics that require semantic…
Configure alerts for latency spikes, error rate increases, and token usage anomalies to catch issues before they impact users.
Log prompts and responses for failed interactions to a secure store for debugging. Implement sampling for successful calls to control storage costs while…
Comprehensive monitoring dashboards enable proactive AI operations management.
Comprehensive observability enables continuous improvement of AI applications.
Building custom observability gives you full control over your data and features. Start simple and add complexity as your needs grow.
Phoenix provides powerful local-first observability that keeps your data private while offering the visualization and analysis capabilities needed for…
Weights & Biases provides a comprehensive platform for LLM observability, evaluation, and collaboration. Its strength lies in combining experiment tracking…
The best tool depends on your specific needs. Start with the simplest option that meets your requirements, and upgrade as your needs grow.
OpenTelemetry provides a standardized way to instrument AI applications, ensuring your observability data is portable across different backends and tools.
Effective tracing reveals the inner workings of AI applications, helping you understand performance bottlenecks, cost drivers, and error sources across your…
Observability transforms AI agents from black boxes into understandable systems. Combine metrics, logs, and traces to gain complete visibility into agent…
Comprehensive audit logging is essential for AI agents in production. It enables debugging, ensures compliance, and provides the visibility needed to build…
Debugging AI applications is different from traditional software. Today I'm exploring tracing and debugging techniques for production AI systems.
AI systems require specialized monitoring beyond traditional application metrics. This guide covers comprehensive observability for production AI.
Advanced techniques for tracing and debugging complex LLM applications in production.
One of my first questions when evaluating any data platform for enterprise use is "how do I see what's actually running?" — and in the early Fabric…
Observability in 2022 moved beyond simple monitoring to true understanding of system behavior. OpenTelemetry became the standard for instrumentation. SLOs…
Azure Monitor Agent provides a unified, secure, and flexible foundation for all monitoring needs.
Azure Managed Grafana provides enterprise observability with seamless Azure integration and reduced operational overhead.
OpenTelemetry provides a future-proof approach to observability. By using vendor-neutral instrumentation with Azure Monitor as the backend, you get the…
Application Insights provides the deep observability needed for modern applications. Combined with Azure Monitor's broader capabilities, it enables…
Azure resources produce several types of diagnostic data: Platform Metrics : Numerical performance data Resource Logs : Detailed operational logs (formerly…
Azure Monitor collects two types of metrics: Platform Metrics : Automatically collected from Azure resources Custom Metrics : Application specific metrics…
Azure Monitor for Prometheus addresses these while maintaining compatibility.
Grafana on Azure became substantially easier when Azure Managed Grafana reached general availability. Before that, teams ran Grafana on AKS or a VM, managed…
Application Insights is the monitoring tool that pays you back in proportion to how much you invest in it. Out of the box you get…
KQL is the query language I've spent the most hours in this year that most developers have never heard of. If you use Log Analytics, Sentinel, Application…
Azure Monitor supports several alert types: Metric alerts Based on metric values crossing thresholds Log alerts Based on Log Analytics query results…
Operational dashboards in Azure used to mean Power BI, a custom data export, and a refresh schedule nobody could remember. Workbooks killed all of that for…
Log Analytics is the database I've spent the most hours in this year that nobody calls a database. It's where every Azure diagnostic, AppInsights trace…