US Federal News Bureau
Written by: Tathagata Sen
Updated 7:51 AM EDT, September 22, 2026

The U.S. Department of Justice (DOJ) is exploring whether commercial AI platforms can handle legal work with less step-by-step human prompting, according to a Nextgov/FCW report.
DOJ posted a market research notice on SAM.gov on September 19, asking vendors for information about AI legal assistant platforms that can analyze documents and conduct research
According to the notice, the market research exercise will not result in a contract award.
The notice describes a potential pilot involving approximately 20 to 30 government users. It would run inside a vendor’s FedRAMP-authorized or commercial cloud environment, using synthetic data or publicly available information rather than real case files.
DOJ wants to see whether a platform can autonomously carry out multi-step legal tasks without an attorney prompting it at each step. Those tasks include document analysis, regulatory review, contract clause extraction, Freedom of Information Act processing, and memo drafting.
The department also wants to know if a platform can process between 5,000 and 100,000 documents in a single session, cite authoritative sources in its legal research, help attorneys draft documents, and integrate with existing document management systems.
To judge performance, DOJ plans to use a mix of usage statistics, government-run quality reviews, vendor timing logs, live demonstrations, checklist reviews, and anonymous staff feedback surveys.
This would build on tools DOJ already uses.
The department has deployed AI features from LexisNexis, Westlaw, and Bloomberg Law for legal research, according to its 2025 AI use case inventory.
DOJ’s approach here is a useful model for chief data officers (CDOs) weighing their own AI rollouts, regardless of industry.
Before committing budget or granting broad access, the department is proposing a structured evaluation with synthetic or publicly available data, a small user group, and multiple independent methods for judging output quality rather than relying on vendor claims.
The decision to pair quantitative measures, like timing logs and usage statistics, with qualitative ones, like government-run quality reviews and staff feedback, would give a fuller picture than relying on any single metric or vendor claim.
Limiting the initial group to 20 to 30 users would ensure that in case of any failure, the impact is contained, while still generating enough real usage to inform a broader rollout decision later. Additionally, because the pilot uses no real case data, a failure would leave DOJ free to investigate and adjust without having to manage the aftermath of a sensitive data incident.
The tasks DOJ wants tested sit closer to agentic AI. Letting a system act across several steps with less step-by-step human prompting raises the stakes on data quality, source reliability, and output verification. This is the kind of rollout where AI governance has to be built in before deployment: knowing what data an agent can touch, requiring it to cite verifiable sources, and having a defined way to catch errors before they reach real casework.