On 2026-08-07, NIST opened a 60-day comment period for the initial public draft of NIST AI 200-2, the TEVV-Athlon Framework. Comments are due by 2026-10-06. The draft is a request for input, not a final standard.
The design change
TEVV means testing, evaluation, verification, and validation. Instead of treating one benchmark as a universal answer, TEVV-Athlon starts from an organisation's evaluation objectives and constructs a context-specific assessment. NIST describes a four-stage method in which Events and Tools produce data about Blocks associated with the concepts being measured.
The framework is intended to be adaptable across statistical machine learning, large language models, multimodal models, and agentic systems. That breadth is a design aim stated by NIST; individual methods still need evidence that they measure the intended property in the intended setting.
Why it matters operationally
A procurement team, product owner, and technical evaluator may all need different evidence from the same system. A useful evaluation plan should therefore record the decision being supported, the relevant environment, failure costs, measurement limits, and acceptance thresholds before tools are selected.
The independent CIRCLE research framework similarly argues for lifecycle and real-world evaluation rather than model scores alone. It is contextual research, not an endorsement or certification of TEVV-Athlon.
Practical next step
Organisations can trial the draft on one bounded use case, preserve raw observations and version identifiers, and document which conclusions the evidence supports. Feedback should address clarity, applicability, missing evaluation activities, and examples while the NIST comment period remains open.