AI In Government4 mins read

JudgeGPT Study: AI Helped Pakistani Judges Clear More Cases, but Training Made the Difference

A field experiment with 1,559 Pakistani judges found JudgeGPT improved case resolution by 6.3 percent when paired with targeted hands-on training, with researchers estimating returns of up to $38.50 per dollar invested.

The core finding: AI improved court productivity when judges were trained

A field experiment with Pakistan’s judiciary found that JudgeGPT, an AI assistant designed for Pakistani trial courts, helped increase case resolution by 6.3 percent. The randomized trial covered 1,559 judges across 118 courts, roughly half of all Pakistani trial court judges. Researchers estimated returns of up to $38.50 per dollar invested, based on the cost of hiring enough judges to match the same output.

The study’s main takeaway is practical: AI tools can improve public-sector productivity, but access alone is not enough. The gains appeared when judges received targeted, hands-on training on what the tool could do, where it could fail, and how to verify its output.

Why training changed the outcome

Judges were split into three groups: one received JudgeGPT access plus targeted training, another received the same AI access with only a general technology-and-law seminar, and a control group attended the seminar without JudgeGPT access. The targeted training included six 90-minute lectures over three weeks after court hours.

AI access alone did little. Judges who received targeted training used JudgeGPT four times as much as those who only attended the general seminar, averaging nearly 60 logins and more than 200 prompts after 40 weeks. The comparison group averaged about 20 logins and fewer than 50 prompts.

Quality held steady, and bias did not increase

The researchers found that ruling quality held steady or improved while judges worked the same hours and reported no change in work-life balance. The appeal rate per 1,000 resolved cases fell slightly. A review of roughly 4,000 judgments found more AI-flagged text, but readability, length, and the number of legal arguments remained stable.

An LLM-based quality check validated by two Pakistani lawyers showed a slight improvement. Rulings from trained judges were rated better in 59 percent of pairwise comparisons, compared with 42 percent in the control group. The study found no evidence that AI use increased gender or religious bias in judicial language.

The safer pattern: support tasks, not automated judging

Chart showing common JudgeGPT tasks and how targeted training shifted use toward editing and summarization
Image credits:Mehmood, Goessmann, Ash (2026)

Anonymized chat logs from about 1,500 judges showed legal research, text editing, and text generation were the most common uses. Around 60 percent of queries sought information about laws, procedures, or legal concepts. Trained judges shifted more toward editing and summarizing, and away from broad legal questions where hallucination risks are higher.

The researchers stressed that the results do not support replacing judges with AI. The practical lesson is to keep final decisions with humans while using AI for bounded support tasks, especially when users are trained to check outputs and understand tool limits.

Discover More