Preliminary data suggest that generic large language models (LLMs) currently lack the precision required for complex statistical tasks, such as power analysis for multisite randomized trials. Project AI4OD will address this gap by establishing a rigorous evaluation benchmark and developing a specialized AI agent designed to democratize high-level statistical expertise. This study and follow-up studies have three aims.
Aim 1: Develop the AI4OD Benchmark Kit for Statistical Reasoning
Aim 2: Evaluate and Profile Generic LLM Performance Disparities
Aim 3: Architect and Prototype the AI4OD Specialized Agent.
Experimental studies investigating moderation and main effects provide the source material for improving the quality of and equity in education by delineating the impact of an intervention, and the contexts, conditions, and sub-populations for which it is most effective. Despite sustained interest in experimental studies, the literature has not developed accessible optimal sampling strategies to help plan powerful and efficient designs to detect these effects. This project, funded by the Spencer Foundation, aims to develop a flexible optimal design framework so that we can design well-powered studies with minimal financial resources. Preliminary results show that the proposed framework can identify much more efficient or powerful designs when contrasted with conventional design frameworks. This project has the potential for broad impacts because it facilitates a fundamental shift in the principles and strategies of study design.
This project will develop the design and analytical frameworks for testing the equivalence of two-group means (or two effects in two studies). It proposes a Monte Carlo confidence interval (MCCI) method to compute the CIs for the tests. In addition, it develops statistical power formulas and tools for the design of studies testing statistical equivalence.