Federal statistics occupy a critical, but often unseen, place in public life. Produced and maintained with taxpayer dollars, these data are used to inform policy, guide business decisions, support research, and help people make daily choices in their lives and understand different things occurring in the country.
Many federal statistical agencies are beginning to adopt artificial intelligence to increase the speed and efficiency of their systems. However, to maintain accuracy and trust in public data, agencies must evaluate AI’s effectiveness, determine what guardrails are needed, and use the technology in ways that are ethical, transparent, replicable, and useful.
The 2026 AI Day for Federal Statistics explored the promise and potential risks of using AI in government data. Convened by the National Academies’ Committee on National Statistics, the Federal Committee on Statistical Methodology, and the National Institute of Statistical Sciences, the event examined active experimentation with AI across several federal agencies including the Census Bureau, Centers for Disease Control and Prevention, Bureau of Economic Analysis, NASA, and the National Center for Health Statistics.
Statisticians in the Driver’s Seat
Statistical experts must be involved early in the implementation and decision-making processes as AI is rolled out at federal agencies, said David Matteson, Director of the National Institute of Statistical Sciences.
“Statisticians must be central to AI leadership,” he asserted. “Statisticians add evaluation, discipline, inference, validity, and accountability for decisions made under uncertainty.”
“Statisticians shouldn’t be downstream reviewers asking to bless the system after [AI tools have] been deployed; they should be partners early and often — leaders, or arm’s length away from leaders, from the beginning.” A major pitfall is modernization without rigorous evaluation, which can create the appearance of progress while introducing new risks.
“Without statistical framing, we risk optimizing the wrong thing extremely well,” Matteson said.
Success in Survey Implementation and Assessment
Surveys are central to federal statistics, but they are expensive, complex, and increasingly difficult to field. AI can reduce manual coding burdens associated with survey data processing while improving the handling of complex text in response fields. Lynda Laughlin, chief of the Industry and Occupations Statistics Branch at the U.S. Census Bureau described how the transition from legacy automated coding systems (autocoders) to large language model-based systems can support better interpretations of industry and job data.
“The legacy system served us well, but it depended heavily on maintenance and performed less well over time as language evolved, and as the quality of responses has gotten worse,” Laughlin explained. “Advances in large language models have significantly expanded what is possible in coding text data.”
While legacy coders are still valuable in data assessment, large language models, or LLMs, can expedite work. In 2012, about 30 percent of cases were autocoded, while 70 percent were sent to the Census Bureau’s National Processing Center for clerical coding. With LLMs, the Census Bureau has been able to code more effectively, resulting in approximately 500,000 fewer cases sent for manual coding each year, Laughlin said.
But this process depends on careful validation. Agencies need to know not only whether a model produces an answer but also where errors concentrate, how those errors affect published statistics, and when human review remains necessary, Laughlin noted.
Kristina Gligorić, assistant professor of computer science at the Johns Hopkins University focused on another promising but delicate use case within surveying — response simulation. From pretesting questions to simulating respondents and improving sampling, LLMs may be able to augment and expedite parts of the survey life cycle.
“Perhaps we could decide who to ask what and in what order,” Gligorić speculated. She continued by raising a central methodical question: How can agencies know whether AI-generated responses provide valid information or simply add noise?
AI systems may also reproduce what researchers expect to hear, a risk sometimes described as social sycophancy.
“A respondent who tells us what we want to hear is not really a great respondent, right?” Gligorić said.
The successful application of LLMs in survey emulation is not whether the simulated response sounds plausible, but whether it improves inference, reduces bias, and can be evaluated against known benchmarks.
Gizem Korkmaz, Vice President of Data Science and AI at Westat, underscored how principles such as validity, reliability, fairness, transparency, and objectivity must be translated into procedures that project teams can mobilize. She highlighted elements to effective governance and implementation, including the need for practical strategies such as human-in-the-loop review, documentation, reproducibility standards, model validation, monitoring, and life-cycle governance.
Institutional Approaches
To facilitate the appropriate adoption of these standards, Zach Whitman, chief AI and data officer at the General Services Administration discussed USAi, a platform intended to provide ready-to-use AI tools, capabilities, and services for federal users. The accessibility of these tools to the research community guarantees a streamlined method to both access federal statistical data, as well as spur adoption for academia, private sector innovators, and other federal agencies, he said.
Benjamin Rogers, acting deputy chief AI officer for CDC, said that the agency’s strategy centers on supporting, strengthening, advancing, and empowering staff use of AI. Its impact across the agency has been substantial, with as much as a 527 percent return on investment, $3.7 million in labor cost savings, and nearly 41,460 CDC staff hours saved and redirected toward higher-value work, he said.
The general public has also used AI tools to self-diagnose, changing the way Americans interact with health care providers.
“Gallup recently put out a survey that noted that as many as 14 million Americans may have canceled a provider visit within the past month because of information that they have received from GenAI tools,” Rogers said. “That’s important to me at the CDC, but we can see the impact across the board.”
Privacy Is Paramount
Privacy and confidentiality should be a major focus when using AI at statistical agencies, speakers emphasized. “We have to think about the fact that AI and machine learning models can retain this information, and [AI] uses it to improve its performance,” said Cordell Golden, data scientist with the CDC’s National Center for Health Statistics.
Lisa Mirel, program director at the National Science Foundation’s National Center for Science and Engineering Statistics addressed privacy enhancing technologies in the National Secure Data Service, emphasizing risk management and the need for a path to implementation focused on privacy and confidentiality. This concern around data permeates into an important early consideration to the production workflow process. Who (or what) selects which data are eligible for a model, how uncertainty is measured, how outputs are reviewed, when results are rejected, and how limitations are communicated to users?
Golden described incorporating AI and machine learning into the data linkage workflow, where improving efficiency and data quality must be balanced with privacy protections and analytic validity.
Continuity
These topics are not new for the National Academies. Previous reports, such as the Toward a 21st Century National Data Infrastructure report series and A Roadmap for Disclosure Avoidance in the Survey of Income and Program Participation, have examined how statistical systems can expand data access and usefulness while protecting individuals and preserving trust. But AI intensifies these concerns because models can make data easier to discover, combine, summarize, and reuse — sometimes in ways that existing governance frameworks did not anticipate, participants said.
Plans Are Worthless, But Planning Is Everything
AI can help the federal statistical system become more efficient and more capable. But to fulfill that potential, agencies must preserve the qualities that make federal statistics valuable in the first place: rigor, impartiality, transparency, confidentiality, and public trust. The future of AI in federal statistics will depend not simply on what models can do, but on how carefully agencies decide what they should do, proceeding deliberately, iteratively, and visibly.
“Accountability, transparency, reproducibility, and public trust are at the forefront of developing these AI tools,” Matteson emphasized.
