
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade
This post is about evaluating a new model against a frozen task set before you change production...
Jul 29, 20263 min read0 reactions0 comments
Tag archive

This post is about evaluating a new model against a frozen task set before you change production...
LLM output prices fell about 94% since 2023: how to cut your AI bill without losing quality...
DeepSeek retires deepseek-chat and deepseek-reasoner on July 24, 2026: your V4 migration...