US Federal News Bureau
Written by: Pritam Bordoloi, Senior Reporter, CDO Magazine
Updated 8:59 PM EDT, August 5, 2026

The National Institute of Standards and Technology (NIST) has launched a new initiative to improve the evaluation of AI models through a secure, sequestered testing environment.
Known as the Artificial Intelligence Technology Evaluation (AITE) program, the effort will enable researchers and developers to assess AI model performance on blind datasets, reducing the risk of train-test data contamination and supporting more rigorous, objective benchmarking.
AITE’s initial focus is on image analysis using large vision-language models (VLMs) across three domains: quantum science, genomics, and public safety.
Additional evaluation tasks are expected to be introduced over time. The program offers two participation tracks. Data providers can submit proprietary datasets and associated tasks, receiving detailed performance assessments of leading AI models on their data. Model providers can submit AI models for testing, gaining insights into how their systems perform across multiple datasets and tasks and how they compare with competing models using standardized metrics.
NIST will provide common datasets, scoring methodologies, and evaluation metrics to improve consistency and comparability across assessments while ensuring evaluation data remains unavailable for model training.
Participation is open to organizations and researchers that agree to AITE’s participation rules and requirements. NIST said the initiative is designed to advance trustworthy AI evaluation and provide developers with clearer measures of model capabilities across real-world applications.