Judge quality with task metrics tied to money or risk, not leaderboard scores. For intake jobs measure precision, recall, and human time per accepted package. For forecasting track error against the method leadership already believes. For assistants measure grounded answer rates and escalation accuracy. Agree on the metric set before coding starts so vendors cannot switch scoreboards after the fact.
Insist on evaluation sets drawn from your domain, frozen during training, and reviewed by people who do the job today. Shadow mode against live work reveals gaps synthetic tests miss. Require logs that show why the system chose an action. Black boxes fail change boards in regulated industries near Newport News. Reproducible runs and versioned artifacts are table stakes for enterprise work.
Process quality counts as much as model quality. Look for weekly demos with real data, written risks, and a willingness to kill a weak approach. Ask how they handle override labeling, rollback, and cost dashboards. Interview the engineers who will staff your project, not only sales. Reference checks should include post-launch support habits, not just launch applause.
We measure ourselves the same way. Baseline first. Publish weekly score movement. Stop when gains plateau until new data or features justify more spend. That discipline protects your capital more than slogan accuracy claims. Bring your metrics into the first conversation so evaluation standards are shared, not imposed late.