AI Models Show Progress in Understanding Supreme Court Complexity, Yet Struggle with Accuracy

Two years after the initial scrutiny of ChatGPT’s capabilities concerning the U.S. Supreme Court, a revisit into AI’s progress reveals notable advances and persistent challenges. The Supreme Court-focused evaluation originally conducted by SCOTUSblog aimed to determine the AI’s accuracy at the onset of its release in 2023. At that time, ChatGPT managed to answer just 42% of the 50 questions correctly.

Since then, the AI models have made strides. The latest assessment involved three contemporary models: 4o, o3-mini, and o1. These iterations improved their performance, yielding 58%, 72%, and 90% accuracy, respectively. The new models have refined their understanding of complex legal concepts and have become more adept at contextual and nuanced analysis. For instance, the AI now accurately identifies historical Supreme Court facts and delves deeper into the intricacies of landmark rulings like Youngstown Sheet & Tube Co. v. Sawyer and Brown v. Allen.

Nonetheless, the AI remains imperfect. Certain models still fabricate or misattribute quotes, as demonstrated when one mistakenly attributed a statement to Chief Justice John Roberts that he never actually made. Despite these fabrications, some models proved reliable against attempts to mislead them with fictitious information.

Moreover, the comparison among the three models showed that while 4o tended to over-extend with unnecessary details and narratives, o1 struck a balance, delivering detailed and accurate answers efficiently. On the other hand, o3-mini faltered in specificity, often providing incomplete or erroneous information.

These advancements suggest that AI, although far from replacing human expertise in legal journalism or research, is evolving. For legal professionals observing the intersection of AI technology and jurisprudence, this evolution might one day offer an auxiliary tool to assist in navigating legal complexities. To explore the detailed exploration of these assessments, visit the SCOTUSblog article.