OpenAI Trains AI Agents on Complex Legal and Business Tasks
Translated & summarized from Behadrei Haredim by baba
OpenAI and Ironclad are collaborating to train AI agents on complex legal and business tasks, with the new GPT-6 Astra model showing significant improvement over its predecessor. The AI agents demonstrated an ability to understand business rules and execute workflows, reducing task completion time from an estimated 37 minutes to 19.2 minutes. However, OpenAI cautioned that these agents cannot yet replace human professionals and that human oversight remains essential for complex operations.
The story in 6 lines · by baba
- OpenAI and Ironclad are researching AI agents' ability to perform complex legal and business tasks.
- The new GPT-6 Astra model achieved a 55% score, a 32% improvement over GPT-5.6 Sol.
- AI task completion time decreased from an estimated 37 minutes to 19.2 minutes with GPT-6 Astra.
- The research involved 11 tasks including NDA creation and purchase approval processes.
- OpenAI stressed that AI agents cannot yet replace human legal professionals.
- Human oversight is still considered essential for complex AI-driven processes.
OpenAI has announced a new research collaboration with Ironclad, a company specializing in contracts, to train and test artificial intelligence agents on complex tasks within the realm of contracts and business processes. The initiative aims to develop AI models that can understand business rules and execute entire professional workflows within dedicated software, moving beyond simple computer operations. The next phase of development focuses on enabling AI agents to grasp how organizations function, maintain context over extended tasks, perform sequential steps, and ensure final outputs adhere to initial requirements.
For this research, OpenAI and Ironclad designed 11 distinct tasks spanning legal, commercial, and procurement domains. These tasks included establishing a non-disclosure agreement, creating a purchase approval process, and updating a legal clause based on user-specified jurisdiction. The companies estimate that an experienced human user would typically take 30 to 40 minutes to complete each task.
The first OpenAI model trained on these tasks, GPT-6 Astra, achieved an average score of 55%, a significant improvement over GPT-5.6 Sol's 41.6%. This represents a 32% increase in average performance. Furthermore, the estimated time to complete these tasks was drastically reduced, with Astra averaging 19.2 minutes per attempt, compared to 37 minutes for GPT-5.6 Sol. An internal model developed during this research reached a score of 63.7%, and OpenAI expressed optimism about incorporating these advancements into future models.
Despite these advancements, OpenAI emphasized that the results do not yet indicate AI agents can replace legal professionals. The company noted that the testing involved only 11 research-specific tasks and that the time savings are simulated rather than proven real-world efficiencies. Human oversight remains crucial for complex processes executed by AI agents. OpenAI is inviting a select group of software companies to join similar research efforts, bringing professional tasks that AI agents currently struggle to perform reliably. The company also confirmed that no private customer contracts from OpenAI or Ironclad were used for training or evaluation in this study.
Mentioned