What does your model evaluation pipeline actually look like in production?
Offline evals, golden sets, LLM-as-judge, human review: what are you running before a model or prompt change ships, and what has caught real regressions?
This is where people who run enterprise technology compare notes before a decision: what they’re evaluating, what they’d pick again and what went wrong. Most members are practitioners. Anyone who works for a vendor carries their company’s name on every post.
The bar is specifics over opinions: what you ran, at what scale and what you learned. Moderators check each new member’s first posts.
Offline evals, golden sets, LLM-as-judge, human review: what are you running before a model or prompt change ships, and what has caught real regressions?
dbt semantic layer, a headless BI layer, or definitions living in each dashboard tool? What has reduced the “two numbers for revenue” problem for you?
Expand-and-contract, online schema change tools, feature flags? Share what has kept large-table migrations boring.
Token spend, retrieval cost, caching, model routing. How are you measuring and reducing cost without hurting quality?
AI engineering, platform, FinOps, security automation? Share what is on your hiring plan.
Account structure, networking, region choice, managed-service lock-in: the lessons that cost the most are the most useful to others.
More PRs, bigger PRs, new review checklists? Share what changed in your process and quality metrics.
Blocking, allow-listing, DLP, enterprise licences, policy and training: what combination is working in practice?
Fleet size, managed vs self-managed, and the tooling that keeps version upgrades from becoming a quarterly fire drill.
East-west bandwidth, GPU cluster fabrics, data movement between sites: what has changed in your network plans?
Renegotiated, moved to another hypervisor, containerised, or went to public cloud? Share your path and what surprised you.
Success criteria, timeboxes, data access, who is in the room. Share what separates useful PoCs from sales theatre.
Inventories, risk tiers, approval workflows, monitoring. What framework did you adopt and what does day-to-day governance look like?
We are hearing a lot about agent pilots and far less about agents still running six months later. What was the first use case, what guardrails did you need, and is it still in production?
Golden paths, templates, portals, self-service environments. Which parts get used every day and which became shelfware?
Commitment management, rightsizing, tagging and showback, architecture changes? Numbers welcome.
Fleet management, updates, security patching and observability at the edge: what tools and practices have held up?
Identity, cloud security posture, AI security, SOC automation, resilience? Share your top two and why.
Vendors are bundling copilots into every suite. How are you deciding what to switch on, and how are you measuring value?
Reserved cloud capacity, neoclouds, on-prem clusters, shared internal pools? What are lead times and utilisation like?
Sampling, tiered retention, OpenTelemetry pipelines, switching vendors? What cut the bill the most?
Identity-first access, device posture, microsegmentation: which parts are done and which keep slipping?
Immutable backups, isolated recovery environments, recovery-time testing. What did your last exercise reveal?
Shared GPU pools and model APIs make showback harder. What allocation model is working?