A.I. Learns to Write in DNA
Scientists prompted A.I. to design and create new viruses for the first time, a milestone for medicine that also raises new biosecurity concerns.
Per a post surfaced on Hacker News, JuliaHub has published a benchmark evaluation comparing OpenAI's GPT-5.6 and Anthropic's Claude Fable 5 on physical AI tasks — a category covering reasoning about real-world environments, robotics, and simulation.
The benchmark's scope goes beyond the text, code, and reasoning tasks that have defined most frontier model comparisons to date. Physical AI evaluation asks models to reason about how objects move, how robots should act, and how simulation environments behave — a meaningfully different test than standard language benchmarks.
The specific results from JuliaHub's evaluation were not available in the materials at the time of this report. The full evaluation lives on JuliaHub's site; the significance of this benchmark will become clearer once those results are reviewed directly.
JuliaHub's choice to test the two latest frontier models on physical tasks signals that this category is becoming a serious axis of model capability measurement. Robotics and embodied AI developers will be watching the results closely.
All comments are reviewed before appearing. Keep it respectful.
Scientists prompted A.I. to design and create new viruses for the first time, a milestone for medicine that also raises new biosecurity concerns.
The Trump administration has no idea on how to handle open-source and open-weight AI models from China.
Google is consolidating AI leadership in California but the reorganization is pushing out the engineers who built its AI division