OpenAI says it has become a C2PA conforming generator, will add Google DeepMind’s SynthID watermarking to OpenAI-generated images, and is previewing a public tool to verify whether an image came from OpenAI.
Archived edition
May 2026 Edition
Issue No. 001 · 55 posts · newest first · last updated May 26, 2026
Notion introduced a Developer Platform with a hosted Workers runtime for custom code, an External Agent API to bring third-party agents into a workspace, and a new `ntn` CLI for developers and coding agents.
A new arXiv paper introduces EngiAI, a LangGraph-based multi-agent reference system, and EngiBench, a benchmark suite to evaluate how LLM agents handle engineering workflows, retrieval, and HPC orchestration.
DeepMind shared interaction principles and demos for an AI-enabled mouse pointer, and says Gemini in Chrome can answer questions about the exact part of a webpage you point to.
DeepMind says it is expanding its Singapore work with new programs in healthcare, education, and sustainability, as part of Google’s national AI partnership with the Singapore Government.
Microsoft Research released MagenticLite plus two small models, MagenticBrain and Fara1.5, aiming to run agentic workflows across the browser and local files on a user’s machine.
NVIDIA describes a “verified agent skills” catalog with scanning, signing, and machine-readable skill cards to help teams trust and audit reusable agent capabilities.
Anthropic’s dashboard says Claude Mythos Preview has generated thousands of vulnerability findings, with 1,596 issues disclosed across 281 open-source projects as of May 22, 2026.
A May 12 arXiv paper proposes GRAFT, mapping tools to special tokens and training on sampled trajectories to improve whether multi-step tool plans follow dependency constraints.
Stability AI released Stable Audio 3.0, including open-weight Small and Medium checkpoints trained on licensed data, plus a Large model offered via its API for higher-volume use.
Cohere released Command A+, an Apache 2.0 open-source MoE model positioned for agentic workflows, multimodal inputs, and long context while targeting self-hosted enterprise deployment.
An arXiv paper reports a “knowing–doing gap” in tool use: models may recognize a tool is needed but still fail to perform the tool call in agent-like workflows.
Google says the Gemini app is adding Daily Brief and Gemini Spark, plus new models like Gemini 3.5 Flash and Gemini Omni, to make it more proactive.
Google DeepMind introduced Co‑Scientist, a multi-agent Gemini-based system for generating and refining scientific hypotheses, and says access will roll out via a research tool.
Anthropic says it acquired Stainless, the company behind its official SDKs, to improve SDK and MCP server tooling for developer experience and agent connectivity.
Anthropic and PwC say they are expanding their alliance, including rolling out Claude Code and Cowork, creating a joint center of excellence, and training 30,000 PwC staff.
A new arXiv paper studies hidden-state trajectories during chain-of-thought and argues you must correct for response length before comparing “reasoning” behavior across tasks.
OpenAI says Codex is now available in preview in the ChatGPT mobile app, letting you check in on long-running work and approve or redirect it from your phone.
Hugging Face and IBM Research introduced an Open Agent Leaderboard to compare how well AI agents handle tool use and multi-step tasks.
Hugging Face and NVIDIA shared a workflow for fine-tuning Cosmos Predict with LoRA and DoRA methods for robot-video generation tasks.
OpenAI and Dell are partnering to support Codex in hybrid and on-premise enterprise environments where companies need tighter data controls.
OpenAI’s AutoScout24 case study shows how one marketplace company is using ChatGPT and Codex to speed engineering work and improve code review.
OpenAI launched DeployCo, a services-style effort meant to help companies turn frontier AI models into working business systems.
OpenAI introduced newer API voice models for realtime conversation, translation, and transcription workflows that developers can build into apps.
OpenAI detailed the Windows sandbox work behind Codex, showing how coding agents can be given useful access without unlimited system permissions.
OpenAI’s Sea Limited case study shows how a large Asian technology company is thinking about Codex and agentic software development.
OpenAI says it is improving how ChatGPT recognizes context in sensitive conversations, especially when risk signals appear over time.
OpenAI says Databricks is using GPT-5.5 for enterprise agent workflows after benchmark gains on office-style knowledge tasks.
Hugging Face and AWS published a practical overview of infrastructure pieces teams use to train, deploy, and serve foundation models.
OpenAI and Malta announced a partnership to give citizens ChatGPT Plus access and training, turning AI adoption into a national digital-skills project.
A new arXiv paper expands MathArena into a continuously maintained evaluation platform for LLM mathematical reasoning, aiming to reduce benchmark saturation and improve comparisons.
Anthropic says it is releasing ten finance agent templates and Claude add-ins for Microsoft 365, so teams can run governed workflows across Excel, PowerPoint, Word, and Outlook.
OpenAI says a TanStack npm compromise impacted two employee devices and it is rotating code-signing certificates, requiring macOS app updates by June 12, 2026.
Meta researchers say tokenization changes scaling behavior and report results suggesting compute-optimal training should track data in bytes, not tokens.
NVIDIA says it is expanding work with ServiceNow on governed autonomous agents, including ServiceNow’s Project Arc and an OpenShell-based runtime for sandboxed, policy-controlled execution.
OpenAI says it is launching the OpenAI Deployment Company and agreeing to acquire Tomoro to bring Forward Deployed Engineers into customer deployments from day one.
Meta researchers introduce NeuralBench and NeuralBench‑EEG, a unified benchmark intended to compare brain-signal AI models across dozens of tasks and many datasets through one framework.
Anthropic says it is handing Petri, its open-source alignment auditing toolbox, to Meridian Labs and releasing Petri 3.0 with more adaptable and realistic behavior tests.
AWS says its MCP Server is generally available, letting AI agents call AWS APIs and read current documentation under IAM guardrails with CloudTrail and CloudWatch visibility.
Google DeepMind says AlphaEvolve, a Gemini-powered coding agent, found algorithm and infrastructure improvements, citing gains in genomics, grid optimization, and systems tuning.
Anthropic describes Model Spec Midtraining (MSM), a training stage that teaches models their behavior spec, and reports large drops in agentic misalignment on scenario tests.
OpenAI says three Realtime API audio models—GPT‑Realtime‑2, GPT‑Realtime‑Translate, and GPT‑Realtime‑Whisper—support voice agents that reason, translate, and transcribe in real time.
OpenAI added an opt-in Advanced Account Security mode that requires passkeys or security keys, tightens recovery, and shortens sessions.
A new arXiv paper introduces AgentFloor, a 30-task tool-use benchmark, and reports many routine agent steps work well on smaller open-weight models.
Google says Gemini API File Search now supports images plus text, metadata filtering, and page citations to ground RAG responses.
OpenAI says GPT‑5.5 Instant, ChatGPT’s default model, is more accurate, cuts hallucinated claims in internal tests, and adds visibility into what context was used for personalization.
NIST’s CAISI says its evaluation of DeepSeek V4 Pro finds the model lags the frontier by about eight months, based on benchmarks spanning cyber, coding, science, reasoning, and math.
NIST’s CAISI signed new agreements with Google DeepMind, Microsoft, and xAI to run pre-deployment evaluations and expand federal research on AI security.
Meta Reality Labs released RL-R CHAT, an egocentric multimodal dataset of group conversations to support hearing-assist and speech enhancement research.
ReasoningBank stores distilled reasoning strategies from both successes and failures, improving tool-using agent performance on web navigation and coding benchmarks.
OpenAI describes a relay-plus-transceiver WebRTC design that keeps voice sessions stable while avoiding huge public UDP port ranges in Kubernetes.
Anthropic researchers report that a small, roughly constant number of poisoned fine-tuning examples can install a backdoor in constitutional classifiers without obvious robustness losses.
Anthropic updated its Responsible Scaling Policy to version 3.2, expanding how its Long-Term Benefit Trust can request and approve external review of risk reports.
OpenAI says AWS customers can access its frontier models, Codex, and Bedrock Managed Agents in limited preview inside existing AWS security and billing workflows.
OpenAI says it surpassed its 10GW by 2029 infrastructure milestone early and is evaluating additional data-center sites to meet rising AI demand.