AIが期間内の動向を整理
OpenAI’s research evaluation and Anthropic’s workbench show AI’s usable scope is expanding
OpenAI released GeneBench-Pro to measure how far AI agents can handle ambiguous judgments in biology research. Anthropic introduced Claude Science as a workbench for scientists. Both moves push AI from a conversation tool toward a practical workplace aid.
SOURCE CHECK
Primary Sources 2
Primary Sources
Key Points
- 1OpenAI released GeneBench-Pro with 129 questions to evaluate judgment tasks in computational biology and clinical genetics.
- 2GPT-5.6 Sol reached up to a 31.5% pass rate, and it is described as helping with work that takes human experts 20-40 hours for a few dollars.
- 3Anthropic’s Claude Science is a workbench that bundles common tools and packages, auditable artifacts, and flexible compute access.
- 4The key takeaway is reduced setup and documentation work, letting teams spend more time on analysis and judgment.
OpenAI released a benchmark for measuring “ambiguous judgments”
GeneBench-Pro is a 129-question evaluation system covering genomics and clinical genetics. It does not just test whether AI can produce an answer; it also examines how it explores data and chooses analysis paths. That makes it important as a practical step toward using AI in research settings.
GPT-5.6 Sol is at the “support” stage, not full replacement
According to the public information, GPT-5.6 Sol reaches up to a 31.5% pass rate and can support work that takes experts 20-40 hours for just a few dollars. This suggests AI can handle at least part of the early-stage research and analysis workflow. It does not mean the final responsibility is replaced.
Anthropic is lowering the setup burden for research environments
Claude Science was introduced as a customizable workbench for scientists. Its features include bundled tools and packages and auditable artifacts. In R&D, the question is not only performance but also whether a tool can reduce preparation and documentation overhead.
Adoption decisions now depend on both “can it work” and “can we keep it”
These two announcements show that AI adoption is no longer determined only by output quality. In real work, the key questions are how much can be delegated, whether results can be audited later, and whether the tool fits existing workflows. Companies evaluating AI should look at both benchmarks and work environments.