What does your model evaluation pipeline actually look like in production?
Offline evals, golden sets, LLM-as-judge, human review: what are you running before a model or prompt change ships, and what has caught real regressions?
Practitioner talk on ai & ml platforms: questions, comparisons and what is working.
2 threads · News in this category
About AI & ML Platforms
Practitioner talk on ai & ml platforms: questions, comparisons and what is working. Practitioners ask, compare and share what worked with AI & ML Platforms. Anyone who works for a vendor is labelled with their company.
Good threads here say what you ran, at what scale and what changed your mind. The most useful answers rise to the top of question threads, and the asker can mark one as accepted. Community guidelines
Offline evals, golden sets, LLM-as-judge, human review: what are you running before a model or prompt change ships, and what has caught real regressions?
Pulse check for FY27 planning conversations.