The compact open-weight model keeps document layouts and fine detail at native resolution, but its capability evidence and deployment readiness remain first-party and incomplete.
Archive
Latest AI intelligence — Page 2
AI reporting, analysis, and product updates across every desk, ordered by publication time.
387 sourced postsThe research caches a teacher model's top token scores and processes KL loss in chunks, avoiding two memory spikes that make long-sequence distillation difficult.
ALTK-Evolve retrieves a task-specific subset of stored lessons instead of sending an entire agent playbook on every step, reducing tokens in IBM's controlled AppWorld runs.
The 3.1-billion-parameter model adds screen understanding, multi-image input, grounding and function calling, while its speed and benchmark claims remain first-party results.
Earth-observation teams can generate geospatial vectors for a chosen place and time, then take the resulting raster into their own analysis tools.
A 2,150-fact benchmark separates whether a model can reveal a fact in familiar context from whether it can produce that fact when directly questioned.
The open reproduction challenge produced thousands of claim-level logbooks, but its own false alarms show why human review still matters.
A January-to-August Hub analysis separates launch excitement from the smaller, older models that remain embedded in real developer workflows.
The Microsoft Research and Xbox prototype lets persistent characters pursue goals, build memories and coordinate while players influence them through conversation.
The research framework scales below the whole-model level and reports lower GPU and power needs on production traces, but is not a generally available service.
The interview study argues that refusal checks and surface-level output tests can miss how chatbot responses affect vulnerable young people in context.
The August 14 cutover leaves existing runs working but moves new custom-model projects toward a public-preview serverless GPU environment.
The August expansion adds partner actions across meetings, travel, entertainment, music and services, with access split by market, account and Gemini mode.
The August 11 snapshot points to heavy voice, visual and image-generation use, but provides no methodology or independent audit for the figures.
The planned signal will change token sampling rather than insert hidden characters, while short, factual, code-heavy or rewritten text may remain difficult to detect.
The company is building a zero-day close and continuous forecasting around approved data, human validation and source-linked outputs, but reports no achieved target.
Developers get more control over long agent runs through subagent tracking, queued commands, headless plan-to-implementation flow and context-preserving app handoff.
The new reasoning-model option reaches five paid plan tiers and eight coding surfaces, with gradual availability, an administrator gate and usage-based billing.
Developers can choose review depth per pull request, while organizations set inherited defaults and account for different resource use.
SL2T 1.0 reaches Pixel 11 with explicit limits against high-stakes use or replacing qualified interpreters.
The builder guide adds deterministic cache breakpoints and same-engine routing hints for repeated agent context, while leaving real-world savings to workload testing.
The new model option spans eight coding environments, but access is gradual, organization administrators must enable a preview policy, and numeric 3.7 pricing is not yet shown in the linked reference.
A randomized simulated-consultation study explores how separate conversation, planning and perception agents can combine in real time, while leaving clinical safety and real-world usefulness unresolved.
LangChain says enterprise customers can keep sensitive agent data in their own AWS accounts while it manages the LangSmith platform lifecycle.