The town square

Feed

Enterprise technology news for the people who buy and run it. Our editors pick stories from company newsrooms and trusted trade press, and every one credits and links its source. You’re seeing Top: recent stories ranked by what readers open, save and discuss. Nobody pays to rank. How ranking works

Narrow it by category below, or follow companies and categories to get their stories in For you and the Monday brief.

Matching stories

Observability

Extend Datadog RUM and Product Analytics to Shopify and Salesforce

Use Datadog RUM and Product Analytics to monitor checkout journeys on Shopify and customer experiences in Salesforce Experience Cloud.

Datadog·via Datadog Blog
Observability

DevRel newsletter — September 2026

Hello from the Elastic DevRel team! In this newsletter, we cover jina-ocr-v1, the latest blogs and videos, and upcoming events like Elastic{ON}.

Elastic·via Elastic Blog
ObservabilityHow-to

What if your agent's hallucinations had a budget? How to start using SLOs for agent behavior

At Grafana Labs, observability is what we do. So as we started building AI agents, we naturally reached for the same instincts we bring to every system: measure it, set targets, and make reliability something you can reason about instead of hope for. That instinct led us somewhere unexpectedly useful. It turns out one of the oldest ideas in reliability engineering, the error budget, maps…

Grafana Labs·via Grafana Labs Blog
Observability

Cribl Stream and Databricks: A faster path to analytics-ready telemetry

The new Databricks Zerobus Destination for Cribl Stream helps organizations move analytics-ready telemetry into Delta Tables with fewer moving parts, less latency, and more control.

Cribl·via Cribl Blog
Observability

Using TypeSafe’s Jev for evals in Datadog Agent Observability

See how we wired Jev into online evals on live spans and offline evals inside Datadog experiments, using a single rubric for both.

Datadog·via Datadog Blog
Observability

Trust, but verify: Atomic claim checking against LLM hallucinations

Hallucinations kept an LLM out of our knowledge base cleanup for years. Splitting generation from atomic claim verification fixed it and cleared a four-year backlog of duplicate articles.

Elastic·via Elastic Blog
Observability

Elastic Cloud on Azure gets a speed boost: Compute-optimized instances on Azure Cobalt ARM

Starting September 24, 2026, Elastic Cloud on Azure adds compute-optimized (ARM) hardware profiles backed by Microsoft's Cobalt ARM processors, delivering up to 37% better throughput than the previous generation Ampere Altra VMs at lower cost.

Elastic·via Elastic Blog
Observability

AI is compressing attack timelines. Cyber resilience needs proof.

How banks can turn the ECB’s AI-cybersecurity expectations into an evidence-driven operating model

Cribl·via Cribl Blog
Observability

Dynatrace achieves ISO/IEC 42001 certification

Dynatrace has achieved the international standard certification for AI management systems, ISO/IEC 42001:2023. The certification applies to the AI management systems (AIMS) supporting the Dynatrace platform, including the AI-powered capabilities Dynatrace develops and embeds in Dynatrace SaaS and Dynatrace Managed. The post Dynatrace achieves ISO/IEC 42001 certification appeared first on…

Dynatrace·via Dynatrace News
Observability

From insight to innovation: How Kiro Crew, Dynatrace, and AWS are helping teams do more

Modern businesses are sitting on a goldmine of information. Your observability data can help you understand what is breaking, what is underperforming, and where your next improvement should come from. The gap isn’t data – it is the distance between seeing something and being able to do something about it That’s exactly what Kiro Crew, […] The post From insight to innovation: How Kiro Crew,…

Dynatrace·via Dynatrace News
Observability

Teaching a 9B model to investigate production alerts

Learn how we fine-tuned Qwen3.5-9B into a specialized agent for change attribution that achieved 87% of GLM-5.3’s Recall@5 roughly 5% of the investigation cost.

Datadog·via Datadog Blog
Observability

Find answers in your logs faster with Datadog’s Tap to Parse

Use Datadog’s Tap to Parse feature to extract searchable fields from unstructured logs across Log Explorer, Log Pipelines, and Observability Pipelines.

Datadog·via Datadog Blog
Observability

Configure RUM SDKs remotely from Datadog

Use RUM Remote Configuration to change SDK sampling rates, privacy settings, and data collection independently of your release cycle.

Datadog·via Datadog Blog
Observability

Elastic Stack 8.19.22 released

Version 8.19.22 of the Elastic Stack was released today. We recommend you upgrade to this latest version . We recommend 8.19.22 over the previous version 8.19.21 For details of the issues that have been fixed and a full list of changes for each product in this version, please refer to the release notes .

Elastic·via Elastic Blog
Observability

The people who help shape Cribl: Celebrating some of our early employees

Cribl·via Cribl Blog
Observability

Is your most expensive talent spending valuable time writing status updates?

Major incidents can turn expert engineers into coordinators and status writers. Shared context can reduce duplicated investigative work. The post Is your most expensive talent spending valuable time writing status updates? appeared first on BigPanda .

BigPanda·via BigPanda Blog
Observability

The 5 Stages of AI-Human Collaboration to Improve Operational Reliability by PagerDuty

AI adoption is now a measurable driver of both uptime and growth. According to PagerDuty’s 2026 State of AI-First Digital Operations report, 75% of organizations... The post The 5 Stages of AI-Human Collaboration to Improve Operational Reliability appeared first on PagerDuty .

PagerDuty·via PagerDuty Blog
Observability

Introducing New Relic Compound Alerts

Discover how New Relic Compound Alerts intelligently correlates related alerts into actionable operational issues to improve incident response and operational efficiency.

New Relic·via New Relic Blog
Observability

Announcing the 2026 Observability Forecast

The Observability Forecast 2026 offers insights from 2,575 IT and engineering leaders and practitioners worldwide on the future of observability.

New Relic·via New Relic Blog
ObservabilityHow-to

How to Build an SRE Agent That Actually Works (Without Blowing the Token Budget)

Learn how to build an SRE agent that delivers accurate, low-latency incident triage with multi-tiered memory and RAG—without blowing your token budget.

New Relic·via New Relic Blog
Observability

Enterprise Service Management Is for More Than IT: HR, Facilities, and Finance Use Cases

09/21/26 What Enterprise Service Management Actually Means Enterprise service management (ESM) is the practice of applying IT service management principles, such as service catalogs, structured intake, automated approvals, and ticket tracking, to departments outside IT. The label sounds abstract until you see it in context: ... The post Enterprise Service Management Is for More Than IT: HR,…

SolarWinds·via SolarWinds Blog
Observability

Grafana Alerting: Scale alert routing without scaling complexity using multiple notification policies

Alert routing often starts simple. A team creates a few contact points, adds some label matchers, and builds a notification policy tree that sends each alert to the right destination. But alerting configurations rarely stay simple. As an organization grows, its notification policy tree must accommodate more teams, services, and routing requirements. Changes for one team still require editing a…

Grafana Labs·via Grafana Labs Blog
Observability

How Adaptive Tail Sampling Works in the OpenTelemetry Collector

A technical deep dive into Honeycomb's adaptive tail sampling processor for the OpenTelemetry Collector: how decisions get made, the samplers available, how thresholds compose with the rest of a sampling pipeline, performance benchmarks, deployment limitations, and how it compares to Refinery.

Honeycomb·via Honeycomb Blog
Observability

Your pod may be requesting 25× more CPU than you think — and Kubernetes won’t tell you

You’ve run kubectl describe pod. Your dashboard shows 40 millicores requested for the api-server container. You probably assumed your monitoring had the full picture. The Kubernetes scheduler actually reserved 1000 millicores — a full CPU core — for it. How could both values be true? If you’re sizing clusters, building cost models, or tuning HPA […] The post Your pod may be requesting 25× more…

Dynatrace·via Dynatrace News
Observability

WebMCP Monitoring: Why Website Uptime Isn’t Enough for AI Agents

AI agents can fail even when your website looks healthy. See how WebMCP monitoring uses synthetic tests to validate the full agent journey, from tool discovery to backend response. The post WebMCP Monitoring: Why Website Uptime Isn’t Enough for AI Agents appeared first on LogicMonitor .

LogicMonitor·via LogicMonitor Blog
Observability

CriblCon 26: Can’t-miss sessions if you want to be a superstar analyst

Cribl·via Cribl Blog
Observability

What TypeSafe’s Jev means for telemetry

Cribl·via Cribl Blog
Observability

Centralized Log Management: A Comprehensive Guide for Engineers

Learn how centralized log management helps developers and engineers improve incident response, reduce costs, and gain better control over log data.

New Relic·via New Relic Blog
Observability

LLM Observability: The 8 Best Tools for Production AI Systems

Compare the top LLM observability tools for production AI — tracing, cost tracking, and evaluation — and see what to weigh before you add another tool to your stack.

New Relic·via New Relic Blog
Observability

Automate Incident Management with PagerDuty Slack by PagerDuty

Most organizations managing major incidents realized that every moment matters. Context-switching between different tools – with multiple web and chat surfaces having to be open... The post Automate Incident Management with PagerDuty Slack appeared first on PagerDuty .

PagerDuty·via PagerDuty Blog
Observability

September 15 SolarWinds Service Desk Release

09/15/26 Discover the new release with insight from our product team: Recap: Service Desk — September 15 New Feature Pre-Release. Newest product updates: Choose the right foundation for each Process Integration Plan Type: All Plans | Status: General Availability Service Desk now gives admins more flexibility ... The post September 15 SolarWinds Service Desk Release appeared first on SolarWinds…

SolarWinds·via SolarWinds Blog
Observability

Dynatrace Release Radar 08.26

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and where to explore them. The post Dynatrace Release Radar 08.26 appeared first on Dynatrace news .

Dynatrace·via Dynatrace News
Observability

Investigate Smarter from Any Agent: Blast Radius and AI Post-Incident Reviews Come to the PagerDuty MCP Server by Antonella Hidalgo Moreno

PagerDuty’s new intelligent Model Context Protocol (MCP) tools, recently announced in Early Access, are a step toward autonomous operations. Here’s how they help customers get... The post Investigate Smarter from Any Agent: Blast Radius and AI Post-Incident Reviews Come to the PagerDuty MCP Server appeared first on PagerDuty .

PagerDuty·via PagerDuty Blog
Observability

AI can speed up development. Can operations keep up?

AI can accelerate software delivery. The operating model must also keep up with the context, ownership, and decisions that follow. The post AI can speed up development. Can operations keep up? appeared first on BigPanda .

BigPanda·via BigPanda Blog
Observability

Red Cards, Comeback Wins, and the Unsung Defense Keeping Modern Business Online

09/15/26 For IT Pro Day 2026, SolarWinds surveyed the global THWACK® community and discovered our IT professionals and sports pros aren’t all that different. Cut through the tech jargon and it’s clear modern IT pros are playing championship-level defense, navigating unforced errors, and quietly keeping ... The post Red Cards, Comeback Wins, and the Unsung Defense Keeping Modern Business Online…

SolarWinds·via SolarWinds Blog
Observability

A Better Way to Monitor Every Digital Journey with LogicMonitor Synthetics and Internet Performance Monitoring

LogicMonitor Synthetics and Internet Performance Monitoring helps ITOps teams catch digital experience issues earlier with outside-in visibility across apps, networks, APIs, and SaaS. The post A Better Way to Monitor Every Digital Journey with LogicMonitor Synthetics and Internet Performance Monitoring appeared first on LogicMonitor .

LogicMonitor·via LogicMonitor Blog
Observability

Digital Experience Monitoring with Grafana Cloud: Session Replay, synthetic checks, and faster investigations

When something breaks in production, the questions that matter most are also the toughest to answer from metrics alone: who was affected, what did they actually see, and is this worth waking someone up for? Answering those questions requires a fuller picture of the issue and its impact on your users. That’s where Digital Experience Monitoring (DEM) in Grafana Cloud comes in. By combining Frontend…

Grafana Labs·via Grafana Labs Blog
Observability

Outage Retrospective: The AWS Outages That Proved the Value of Independent Monitoring

Internet Performance Monitoring detected the AWS outages before AWS acknowledged them. Learn why independent, outside-in monitoring is critical to faster incident response. The post Outage Retrospective: The AWS Outages That Proved the Value of Independent Monitoring appeared first on LogicMonitor .

LogicMonitor·via LogicMonitor Blog
Observability

Why Government IT Needs AI-Powered Observability

Legacy systems, cloud growth, and AI are reshaping government IT. See how connected visibility helps teams respond faster and keep essential services running. The post Why Government IT Needs AI-Powered Observability appeared first on LogicMonitor .

LogicMonitor·via LogicMonitor Blog
ObservabilityHow-to

The AI Economy Has a Senior Engineer Problem. Here’s How to Solve It by PagerDuty

According to a 2025 report from Ravio, entry-level hiring (especially in engineering roles) has collapsed by more than 73% due to increasing AI capabilities. That... The post The AI Economy Has a Senior Engineer Problem. Here’s How to Solve It appeared first on PagerDuty .

PagerDuty·via PagerDuty Blog
Observability

How Canvas Powers the AI Agent Development Feedback Loop

AI agents need a real feedback loop, not just reactive debugging. This post walks through how Honeycomb's Canvas powers each stage—instrumenting agents with OpenTelemetry, understanding a single run, finding the problems worth fixing, shipping the fix, and proving it worked—plus tips for getting the most out of Canvas.

Honeycomb·via Honeycomb Blog
Observability

Investigate production issues from Cursor with the Dynatrace plugin

Dynatrace is now listed in the Cursor Marketplace as a verified plugin. One install connects Cursor to your Dynatrace environment through the Model Context Protocol (MCP) server and loads 30 Dynatrace skills, giving Cursor access to live Dynatrace data and the context to effectively analyze the data. Connecting your environment Connect Cursor to your Dynatrace […] The post Investigate production…

Dynatrace·via Dynatrace News
Observability

The enterprise changed. ITOps didn’t.

The enterprise modernized its technology stack. Its operating model still depends on human-speed work. The post The enterprise changed. ITOps didn’t. appeared first on BigPanda .

BigPanda·via BigPanda Blog
Observability

What’s new from BigPanda: September 2026 Product Updates

A focus on the BigPanda platform value and foundational updates to AI Detection and Response Most teams we talk to are fighting the same battle. The knowledge needed to make a decision already exists somewhere. However, it’s locked in a tool your team isn’t looking at, or in the head of the one engineer who’s […] The post What’s new from BigPanda: September 2026 Product Updates appeared first on…

BigPanda·via BigPanda Blog
Observability

Custom labels in Grafana Cloud Synthetic Monitoring: New updates for consistency and ease-of-use

Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in Synthetic Monitoring have worked a little differently: they only lived on a single sm_check_info metric, and Grafana Cloud prefixed each one with label_ . To make custom labels in Synthetic Monitoring…

Grafana Labs·via Grafana Labs Blog
Observability

Observability ROI: Real Savings From Real Deployments

Observability can cut alert noise, speed up incident response, reduce downtime, and give engineers more time for planned work. LogicMonitor customers have used those gains to lower costs and make better infrastructure decisions. The post Observability ROI: Real Savings From Real Deployments appeared first on LogicMonitor .

LogicMonitor·via LogicMonitor Blog
ObservabilityHow-to

How to monitor Cypress tests with Grafana Cloud

If your Cypress suite has tests that fail more often or run slower, you know it can be hard to figure out the pattern from a single job. It could be one spec that slowed down, or a single test that fails, or maybe the entire suite is trending slower. The root cause could be a bug in the app, or a flaky test, or something else. Your terminal output and CI log will tell you what happened on a…

Grafana Labs·via Grafana Labs Blog
Observability

AI Norms & Values, Part 3 of 3: Things We Hold True

The final part of Honeycomb's AI Norms & Values series: the principles the company holds true about AI as a tool, ownership of work, and rising standards; how it actually uses AI day to day; usage patterns for respecting each other's time; and where it stands on AI's ethical externalities like energy use, IP, bias, and wages.

Honeycomb·via Honeycomb Blog
ObservabilityComparison

Wide Events vs. Three Pillars: AI Observability Costs

AI agents make telemetry costs harder to predict. This post compares the three pillars against the wide event model, and explains why wide events keep AI observability costs predictable without sacrificing the context engineers need.

Honeycomb·via Honeycomb Blog
Observability

Relational Query Superpowers

See how Honeycomb's relational query keywords—root, parent, child, any, any2, any3, and none—let you pull attributes from anywhere in a single trace into one query, walked through with a real checkout-error investigation.

Honeycomb·via Honeycomb Blog
Observability

Introducing Swarm Investigation from BigPanda: Autonomous, multi-agent IT incident investigation

BigPanda introduces Swarm Investigation, the latest capability for the BigPanda AI Incident Assistant that adds autonomous, multi-agent investigation. The post Introducing Swarm Investigation from BigPanda: Autonomous, multi-agent IT incident investigation appeared first on BigPanda .

BigPanda·via BigPanda Blog
Observability

Automate Product Analytics reports with your agent and the CX CLI

Every page view, click, and session your RUM SDK captures lands in Coralogix as a log event under the cx_rum subsystem — the raw data behind how people actually use your product. You can turn it into a shareable report without writing a single query. Just ask your coding agent. Your agent queries that data […] The post Automate Product Analytics reports with your agent and the CX CLI appeared…

Coralogix·via Coralogix Blog
Observability

See It, Approve It, Revoke It: Scoped OAuth for Public Apps by Aatharsha Jeyachelvan

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about... The post See It, Approve It, Revoke It: Scoped OAuth for Public Apps appeared first on PagerDuty .

PagerDuty·via PagerDuty Blog
Observability

We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents.

How Checkly rewrote a Node.js service handling 92M messages a day into Go using AI agents, and the test harness that made it safe to ship.

Checkly·via Checkly Blog
Observability

Loop Engineering Guardrails for iGaming with Claude Code and CX CLI

A guardrail is a policy that sits between the user and the model, scoring every prompt and response against rules the business has written and blocking the ones that break them. For an agent the stakes are higher than a bad answer, because an agent holds tools: the same prompt that would produce an awkward […] The post Loop Engineering Guardrails for iGaming with Claude Code and CX CLI appeared…

Coralogix·via Coralogix Blog
Observability

How iGaming operators connect revenue metrics to the operational telemetry

Global iGaming gross gaming revenue reached $115 billion in 2026, twelve percent up on the year before. A growing number of iGaming operators run on Coralogix, and what they use it for has moved well past uptime. Soft2Bet powers ninety percent of its internal dashboards from Coralogix and treats it as the single source of […] The post How iGaming operators connect revenue metrics to the…

Coralogix·via Coralogix Blog
Observability

What Is the Parquet File Format? Complete Guide

When your analytics bill depends on bytes scanned, your storage layout can determine whether you read a few relevant chunks or an entire dataset. Apache Parquet is an open source column-oriented format built for efficient storage and retrieval, and it groups values by column so your query engine can read only the fields a query […] The post What Is the Parquet File Format? Complete Guide appeared…

Coralogix·via Coralogix Blog
Observability

Zero-Code Instrumentation in Kubernetes Without the Instrumentation CRD

The OpenTelemetry Operator changed how teams approach telemetry collection in Kubernetes. The core appeal of zero-code instrumentation is that you can bring up telemetry inside application containers to collect traces, metrics, and logs without touching your source code or rebuilding your container images. However, if you follow the default OpenTelemetry Operator documentation, you quickly run…

Coralogix·via Coralogix Blog
ObservabilityWhy it matters

Metabase Security Incident

On 3 August 2026, an attacker exploited a zero-day vulnerability in Metabase, the third-party analytics tool we use internally, and gained read access to a database holding a copy of Checkly operational data. The attacker bypassed authentication and obtained an administrator session on our Metabase Cloud instance. Metabase has since blocked the attack, patched the vulnerability, and published a…

Checkly·via Checkly Blog
ObservabilityWhy it matters

Elastic 9.5: Columnar, VectorDB index mode & auto-calibration, and AI-driven alert triage

Today, we are pleased to announce the GA of Elastic 9.5 as the latest version of the Elasticsearch Platform. This release includes a range of new features, including Columnar Mode, VectorDB index mode, and Agent Builder enhancements. Learn more.

Elastic·via Elastic Blog
Observability

CLIs are more token-efficient than MCP. Or are they?

MCP servers used to be the token-hungry option. Deferred tool loading and resource links changed that. A side-by-side Playwright MCP vs CLI token comparison.

Checkly·via Checkly Blog
Observability

How Upstash Monitors Every Redis Replica with Checkly

How Upstash monitors uptime for every Redis replica: Checkly URL monitors in Terraform, checks from 18 global locations, and an end-to-end QStash check.

Checkly·via Checkly Blog
Observability

Autoscaling Checkly Private Location Agents in Kubernetes with KEDA

Monitor your Private Location health with Checkly, and autoscale agents with KEDA on live load. Stop guessing at agent capacity.

Checkly·via Checkly Blog
Observability

The High-Performance DBA: Breaking the Burnout Cycle in Modern Database Teams

04/01/26 In today’s data-driven environments, the role of the database administrator (DBA) has never been more critical or more demanding. Modern database administration now spans on-premises data centers, Microsoft SQL Server, Oracle, and AWS-hosted databases, as well as third-party managed services. As organizations scale across ... The post The High-Performance DBA: Breaking the Burnout Cycle…

SolarWinds·via SolarWinds Blog
Observability

Inside the Black Box: Bridging the Database Observability Gap

03/27/26 Over the past 15 years, Agile and DevOps have transformed how organizations deliver software. Agile emphasizes iterative development, collaboration, and continuous feedback, while DevOps brings development and operations together to automate and accelerate software delivery. Teams now rely on rapid iterations, shared responsibility, and ... The post Inside the Black Box: Bridging the…

SolarWinds·via SolarWinds Blog