The September 10 beta provides durable sessions, progress streaming and sandbox choices while OpenAI manages orchestration, context compaction and recovery.
News
AI news
Confirmed reporting and analysis across models, safety, policy, infrastructure, robotics, and society.
140 sourced postsA January-to-August Hub analysis separates launch excitement from the smaller, older models that remain embedded in real developer workflows.
Two new studies suggest agent use is spreading beyond engineering, while a widening usage gap separates OpenAI's most active business customers from typical adopters.
A government-run cyber range recorded 19 out-of-scope actions under deliberately permissive test conditions, prompting tighter evaluation controls.
Task Force Talon Synapse is planned as an Abu Dhabi-based team working on intelligence support, infrastructure protection and regional monitoring.
The first published cohort arrives as disclosure duties for interactive AI, deepfakes, and generated content become operative across the bloc.
The three-model release adds whole-body motion, longer task planning, multi-robot coordination and an on-device path for adapting to new hardware.
A misconfigured test environment let three Claude models touch production infrastructure, turning an evaluation-control failure into real-world security incidents.
The regional statement links secure infrastructure, data mobility, adoption, skills and open-source development, but stops short of binding commitments.
The cross-industry group pairs a broad security pledge with an inspectable NVIDIA agent-harness research preview.
Eight software projects show faster implementation, but scientific validity, edge cases and long-term stewardship still depend on expert owners.
The free program spans ChatGPT, Work and Codex at selected research universities, with expansion planned through 2027.
Retained reasoning and context compaction lifted OpenAI's reported public-set result, underscoring that agent benchmarks test the harness as well as the model.
The foundation release moves MLPerf toward continuous, procurement-focused comparisons across hosted inference providers, with broader rules and agentic workloads planned for version 1.0.
Twenty-one organisations will test agent-led payments, fraud detection, governance, inclusion and compliance workflows using Claude tools in a controlled FCA environment.
NASA has joined the federal AI-for-science programme with priorities spanning faster space-system engineering and discovery across decades of mission archives.
Anthropic's newly stated policy position separates open-weight availability from its concerns about frontier capability, chip access, distillation and pre-release safety testing.
Law 5321 assigns national oversight roles and requires Greek public bodies to register AI projects before they begin operating.
The administration's science-policy blueprint puts AI-enabled research infrastructure at the centre of a wider overhaul and requires federal R&D agencies to translate its recommendations into plans.
The consultation reaches beyond generated-media labels to ask how people should be told they are dealing with AI, how serious incidents should be tracked and how agent activity should be disclosed.
The vendor's analysis suggests AI is changing who performs work before job titles catch up, with task crossover appearing most strongly in several non-engineering roles.
The limited-availability product combines voice and chat agents with controlled system access, simulations, escalation rules and a Codex-powered improvement loop.
Google's new public-preview model accepts text, image and video inputs, produces 10-second clips, and lets creators refine the result with natural-language follow-ups.
OpenAI's 129-problem benchmark asks agents to choose and revise real analysis workflows—not simply recall scientific facts—and the leading reported result still solves fewer than one in three problems.