Datadog vs Grafana stack vs OpenTelemetry + open source
If you have run more than one, compare them on cost at your scale, time to value and on-call experience.
Evaluating Observability this quarter? This page brings together what changed, which companies are most active, what practitioners are asking and how to build a shortlist.
Logs, metrics, traces and synthetic monitoring, plus the shift to OpenTelemetry, cost-aware telemetry and AI-assisted incident response.
Start with the news. These are the Observability stories getting the most attention right now, so you know what has shifted before you talk to anyone.
Use Datadog RUM and Product Analytics to monitor checkout journeys on Shopify and customer experiences in Salesforce Experience Cloud.
Hello from the Elastic DevRel team! In this newsletter, we cover jina-ocr-v1, the latest blogs and videos, and upcoming events like Elastic{ON}.
At Grafana Labs, observability is what we do. So as we started building AI agents, we naturally reached for the same instincts we bring to every system: measure it, set targets, and make reliability something you can reason about instead of hope for. That instinct led us somewhere unexpectedly useful. It turns out one of the oldest ideas in reliability engineering, the error budget, maps…
The new Databricks Zerobus Destination for Cribl Stream helps organizations move analytics-ready telemetry into Delta Tables with fewer moving parts, less latency, and more control.
See how we wired Jev into online evals on live spans and offline evals inside Datadog experiments, using a single rubric for both.
Hallucinations kept an LLM out of our knowledge base cleanup for years. Splitting generation from atomic claim verification fixed it and cleared a four-year backlog of duplicate articles.
Starting September 24, 2026, Elastic Cloud on Azure adds compute-optimized (ARM) hardware profiles backed by Microsoft's Cobalt ARM processors, delivering up to 37% better throughput than the previous generation Ampere Altra VMs at lower cost.
How banks can turn the ECB’s AI-cybersecurity expectations into an evidence-driven operating model
Those stories keep returning to the same names. These are the companies we track in this category; each profile has its news, solutions and published reviews.
Datadog is an American observability and security company founded in 2010 by Olivier Pomel and Alexis Lê-Quôc and headquartered in New York City. Its SaaS platform provides infrastructure monitoring, application performance monitoring, log management, real user monitoring, synthetic testing, cloud security and incident management for cloud-scale applications. Datadog has been listed on Nasdaq (DDOG) since 2019.
Dynatrace is an observability and application security company founded in Linz, Austria, in 2005 and headquartered in Waltham, Massachusetts. Its AI-powered platform covers infrastructure and application monitoring, log analytics, digital experience monitoring, application security and automation, built on its Grail data lakehouse and Davis AI engine. Dynatrace has been listed on the NYSE (DT) since 2019.
Splunk is a US software company, founded in 2003 in San Francisco, that makes data platforms for security and observability. Its products include Splunk Enterprise, Splunk Cloud Platform, Splunk Enterprise Security (SIEM), Splunk SOAR, Splunk Observability Cloud and IT Service Intelligence. Cisco completed its acquisition of Splunk in March 2024, and Splunk now operates as part of Cisco.
New Relic is an observability company that provides a SaaS platform for monitoring applications, infrastructure, logs, browsers, mobile apps and AI workloads. Its all-in-one platform combines telemetry data with APM, distributed tracing, dashboards and alerting under usage-based pricing. Founded in 2008 by Lew Cirne, New Relic is headquartered in San Francisco and was taken private by Francisco Partners and TPG in 2023.
Grafana Labs is the company behind Grafana, the open source visualisation and dashboarding tool, and the LGTM observability stack of Loki, Grafana, Tempo and Mimir. It sells Grafana Cloud and Grafana Enterprise for metrics, logs, traces, profiles, load testing with k6 and incident response. Founded in 2014 and operating remote-first, Grafana Labs is privately held.
Elastic is the company behind Elasticsearch, Kibana, Logstash and Beats, collectively known as the Elastic Stack. Founded in 2012 by Shay Banon and others, it is incorporated in the Netherlands and operates as a distributed company with a US headquarters in San Francisco. It offers search, observability and security solutions, delivered via Elastic Cloud or self-managed. Elastic is listed on the NYSE (ESTC).
Honeycomb is a US observability company whose platform lets engineering teams query high-cardinality, high-dimensionality event and trace data to debug production systems. It is built around OpenTelemetry and offers distributed tracing, service-level objectives, BubbleUp anomaly analysis and log analysis. Honeycomb was founded in 2016 by Charity Majors and Christine Yen, is based in San Francisco and is privately held.
Chronosphere is a cloud native observability company founded in 2019 by Martin Mao and Rob Skillington, creators of the M3 metrics platform at Uber. It offers the Chronosphere Observability Platform for metrics, traces, logs and events, with controls to shape and reduce telemetry data volumes, and Calyptia-based telemetry pipelines. Palo Alto Networks completed its acquisition of Chronosphere in January 2026.
ServiceNow is a US enterprise software company headquartered in Santa Clara, California, founded in 2004 by Fred Luddy. Its ServiceNow AI Platform automates workflows across IT service management, IT operations, security, customer service, HR and field service, with AI agents, low-code app building and observability capabilities. ServiceNow is listed on the New York Stock Exchange.
IBM is a US technology company providing hybrid cloud and AI software, infrastructure and consulting. Its portfolio includes Red Hat OpenShift, the watsonx AI and data platform, IBM Z mainframes and Power servers, IBM Storage, Db2 databases, and automation and observability software such as Instana and HashiCorp Terraform. Founded in 1911 and headquartered in Armonk, New York, IBM is listed on the NYSE.
Sumo Logic is a US cloud-native log analytics company founded in 2010 and headquartered in Redwood City, California. It provides a SaaS platform for log management, observability and security analytics, including cloud SIEM, used by DevOps, IT and security teams. Sumo Logic was listed on Nasdaq from 2020 until Francisco Partners took it private in 2023.
LogicMonitor is a SaaS observability company whose LM Envision platform monitors hybrid IT infrastructure, including networks, servers, cloud services, containers, applications and logs, with AI-driven event correlation through its Edwin AI product. It serves enterprise IT and operations teams and managed service providers. Founded in 2007 and headquartered in Santa Barbara, California, it is privately held and backed by Vista Equity Partners.
SolarWinds is an IT management and observability software company headquartered in Austin, Texas, founded in 1999. It sells SolarWinds Observability for SaaS and self-hosted environments, Network Performance Monitor, Server & Application Monitor, Database Performance Analyzer and service desk software. In 2025 SolarWinds was taken private by private-equity firm Turn/River Capital.
Coralogix is an observability platform company founded in 2014 by Ariel Assaraf. It provides a SaaS platform for logs, metrics, traces, real user monitoring and security data, using an in-stream analysis architecture that processes telemetry before indexing to help control data storage costs. Coralogix also offers AI observability capabilities for monitoring generative AI applications. The company is privately held.
PagerDuty is a digital operations management company whose PagerDuty Operations Cloud provides incident management, on-call scheduling and alerting, AIOps event correlation, runbook and process automation, and customer service operations. It serves engineering, IT operations and support teams responding to incidents in critical systems. Founded in 2009 and headquartered in San Francisco, PagerDuty is listed on the New York Stock Exchange.
BigPanda is a US AIOps company founded in 2012. Its platform uses machine learning and AI to correlate alerts from monitoring, observability and change tools into incidents, surface probable root causes and automate incident response workflows, helping IT operations and site reliability teams in large enterprises reduce alert noise and resolve outages faster.
Kentik is a US network observability company whose SaaS platform analyses flow data, device metrics, cloud flow logs, routing and synthetic test results to help teams monitor, troubleshoot and secure networks. It serves enterprises, cloud-native companies and service providers, with traffic analytics, DDoS detection, cloud network visibility and capacity planning. Founded in 2014, Kentik was acquired by Infoblox in August 2026.
Checkly is a synthetic monitoring and API monitoring company founded in 2018 and based in Berlin, Germany. Its platform uses Playwright-based browser checks and API checks defined as code, with a CLI and Terraform provider, to test and monitor web applications and APIs in production. It targets developers and SRE teams practising monitoring as code.
Cribl is a telemetry data management company founded in 2017 and headquartered in San Francisco. Its products, including Cribl Stream, Cribl Edge, Cribl Search and Cribl Lake, let IT and security teams collect, route, reduce, enrich and search logs, metrics and traces across observability and SIEM tools. Cribl was founded by Clint Sharp, Dritan Bitincka and Ledion Bitincka and is privately held.
Microsoft is a global technology company whose enterprise portfolio spans the Azure cloud platform, Microsoft 365 productivity software, Dynamics 365 business applications, security products and developer tools such as GitHub. Founded in 1975 and headquartered in Redmond, Washington, Microsoft is one of the largest providers of cloud and AI services to businesses.
Vendor news only tells you half of it. This is what people working in Observability are discussing and asking each other.
If you have run more than one, compare them on cost at your scale, time to value and on-call experience.
Sampling, tiered retention, OpenTelemetry pipelines, switching vendors? What cut the bill the most?
With the news, the companies and the open questions in view, pick three to five companies whose products match your problem, read their recent stories and put your remaining questions to the community before you speak to sales.
If you want more background, here is more of the recent coverage.
Dynatrace has achieved the international standard certification for AI management systems, ISO/IEC 42001:2023. The certification applies to the AI management systems (AIMS) supporting the Dynatrace platform, including the AI-powered capabilities Dynatrace develops and embeds in Dynatrace SaaS and Dynatrace Managed. The post Dynatrace achieves ISO/IEC 42001 certification appeared first on…
Modern businesses are sitting on a goldmine of information. Your observability data can help you understand what is breaking, what is underperforming, and where your next improvement should come from. The gap isn’t data – it is the distance between seeing something and being able to do something about it That’s exactly what Kiro Crew, […] The post From insight to innovation: How Kiro Crew,…
Learn how we fine-tuned Qwen3.5-9B into a specialized agent for change attribution that achieved 87% of GLM-5.3’s Recall@5 roughly 5% of the investigation cost.
Use Datadog’s Tap to Parse feature to extract searchable fields from unstructured logs across Log Explorer, Log Pipelines, and Observability Pipelines.
Use RUM Remote Configuration to change SDK sampling rates, privacy settings, and data collection independently of your release cycle.
Version 8.19.22 of the Elastic Stack was released today. We recommend you upgrade to this latest version . We recommend 8.19.22 over the previous version 8.19.21 For details of the issues that have been fixed and a full list of changes for each product in this version, please refer to the release notes .
Major incidents can turn expert engineers into coordinators and status writers. Shared context can reduce duplicated investigative work. The post Is your most expensive talent spending valuable time writing status updates? appeared first on BigPanda .
AI adoption is now a measurable driver of both uptime and growth. According to PagerDuty’s 2026 State of AI-First Digital Operations report, 75% of organizations... The post The 5 Stages of AI-Human Collaboration to Improve Operational Reliability appeared first on PagerDuty .
Discover how New Relic Compound Alerts intelligently correlates related alerts into actionable operational issues to improve incident response and operational efficiency.
The Observability Forecast 2026 offers insights from 2,575 IT and engineering leaders and practitioners worldwide on the future of observability.
Learn how to build an SRE agent that delivers accurate, low-latency incident triage with multi-tiered memory and RAG—without blowing your token budget.
09/21/26 What Enterprise Service Management Actually Means Enterprise service management (ESM) is the practice of applying IT service management principles, such as service catalogs, structured intake, automated approvals, and ticket tracking, to departments outside IT. The label sounds abstract until you see it in context: ... The post Enterprise Service Management Is for More Than IT: HR,…
Alert routing often starts simple. A team creates a few contact points, adds some label matchers, and builds a notification policy tree that sends each alert to the right destination. But alerting configurations rarely stay simple. As an organization grows, its notification policy tree must accommodate more teams, services, and routing requirements. Changes for one team still require editing a…
A technical deep dive into Honeycomb's adaptive tail sampling processor for the OpenTelemetry Collector: how decisions get made, the samplers available, how thresholds compose with the rest of a sampling pipeline, performance benchmarks, deployment limitations, and how it compares to Refinery.
You’ve run kubectl describe pod. Your dashboard shows 40 millicores requested for the api-server container. You probably assumed your monitoring had the full picture. The Kubernetes scheduler actually reserved 1000 millicores — a full CPU core — for it. How could both values be true? If you’re sizing clusters, building cost models, or tuning HPA […] The post Your pod may be requesting 25× more…