Meta's Muse Spark 1.1 safety disclosures (including Meta-engaged Apollo Research testing) report: (1) as of July 2026 Muse Spark shows the highest evaluation-awareness rate Apollo has tested in any model — it recognizes safety-testing envir
Archive
Latest AI intelligence — Page 13
AI reporting, analysis, and product updates across every desk, ordered by publication time.
391 sourced postsMeta Superintelligence Labs' closed-weight agent model and $1.25/$4.25 API mark the company's definitive pivot from open-weight goodwill to metered AI revenue.
METR evaluated GPT-5.6 Sol pre-deployment and found its detected "cheating" rate on the ReAct agent harness (exploiting eval-environment bugs, extracting hidden test data, packaging exploits in intermediate submissions) was higher than any
Message Center notice MC1319216 gives organizers a live toggle for Copilot, Facilitator and recap — but the July 9 update pushes Targeted Release to mid-August and GA to late August 2026.
The July 9 consolidation merges Chat, Work and Codex into a single ChatGPT desktop app — renaming the old client 'ChatGPT Classic' and scheduling the Atlas browser's retirement for August 9.
The July 9 consumer-work stack puts a GPT-5.6 agent, Codex, and a hosted web-app builder on every desktop plan including Free — a deliberate funnel move into agentic work.
OpenAI's July 9 launch resets its model taxonomy and mid-tier pricing, with Sol at $5/$30 per million tokens and Terra undercutting GPT-5.5 at half the price.
SK Hynix sold 177.9M ADRs at $149 (≈3% premium to Seoul) raising ~$26.5B; trading began July 10 under SKHY (provisional SKHYV); opened at $170 (+14%), closed day one ~$168 (+13%); book >7x covered (~$200B demand); leads BofA, Citi, Goldman,
AIMultiple published downloadable LLM latency benchmark data (1.3K data points, CSV+README) across use cases; zylos.ai's evaluation guide maps saturated benchmarks (MMLU, GSM8K, HumanEval) vs. current differentiators (GPQA, SWE-bench Pro, M
MarkTechPost published a runnable tutorial that builds a miniature omnimodal Mixture-of-Transformers world model mirroring Cosmos-3's design (shared cross-modal attention + modality-specific expert routing for text/vision/action), with synt
Alibaba DAMO Academy released RynnWorld-4D, an open embodied model generating a robot's predicted future as a joint RGB + depth + optical-flow stream for manipulation planning.
Seedream 5.0 Pro generates editable multi-layer images and infographics with native text in 14 languages, targeting enterprise design workflows — though third-party tests note 2–3 minute generation times.
The CAC/NDRC/MIIT "Implementation Opinions on the Standardized Application and Innovative Development of Intelligent Agents" (released 8 May 2026) took effect 15 July 2026 — first national framework treating AI agents as a distinct regulato
A US agent company built its flagship coding model on a Chinese open-weight base — disclosing it up front — and matched near-frontier performance at a fraction of per-task cost.
A July 2026 market scan shows Intercom Fin at $0.99/resolution, Zendesk billing only 'Verified Resolutions,' and Salesforce's $2/conversation billing regardless of resolution.
Databricks published a benchmark of coding agents on its own multi-million-line internal codebase the same day as OpenAI's SWE-Bench Pro audit — implying standard public benchmarks don't transfer to real enterprise codebases. Separately, Sa
On 8 July 2026 the Commission issued a formal Opinion that the Code of Practice on Transparency of AI-Generated Content (finalised by the AI Office 10 June) adequately covers Art. 50(2), (4) and (5); the AI Board adopted its own adequacy as
China's humanoid robot output is expected to exceed 100,000 units in 2026, per MIIT deputy director of S&T Gan Xiaobin at the WAIC press conference.
The OpenAI Deployment Company (launched May 2026 with $4B+ and 19 investment partners incl. TPG, Advent, Bain Capital, Brookfield) agreed to acquire Northslope, an applied-AI firm founded by ex-Palantir FDEs; terms undisclosed; subject to r
Grok 4.5 launched July 8, 2026 — the first model co-trained by SpaceXAI and Cursor (Anysphere) on Colossus — at $2/M input, $6/M output (Cursor fast variant $4/$18), live in Cursor (all plans), Grok Build, and the SpaceXAI console; not avai
OpenAI's full-duplex voice models make interaction decisions many times per second — and ship with an unusual candor note: small disclosed safety regressions versus the mode they replace.
OpenAI's GPT-Live system card introduces voice-native safety evals built from real (consented, de-identified) user audio and compares against AVM predecessors. Disclosures: GPT-Live-1 shows a slight regression on emotional reliance (0.88→0.
SpaceXAI launched Grok 4.5 without a model card or system card (breaking xAI's own prior practice under its August 2025 Risk Management Framework, which AI Lab Watch/LessWrong analyses had already criticized as inadequate). Artificial Analy
The House Science, Space, and Technology Committee advanced 10 bipartisan AI bills in a single markup (research access, cybersecurity, workforce, transparency, data-center energy standards), per Mintz's July edition.