❌

Reading view

AWS Weekly Roundup: BYOM for Amazon RDS for SQL Server, AWS IoT Device SDK for Swift, and more (June 8, 2026)

This week, the AWS IoT Device SDK for Swift reached general availability. As a member of the Swift Server Workgroup (SSWG), this one caught my attention. The SDK brings production-ready MQTT 5 connectivity, Device Shadow, Jobs, and fleet provisioning to Swift developers on macOS, iOS, tvOS, and Linux.

Swift on IoT and Edge devices, an AI generated illustration

I’m curious to see what you will build with it. Swift on the server has matured over the past few years, and now it reaches IoT devices too. This connects to a broader trend of running Swift at the edge. WendyOS, for example, is an open-source operating system for physical AI that offers first-class Swift support for deploying apps to NVIDIA Jetson and Raspberry Pi hardware. Between server-side Swift, IoT, and edge computing, the language is showing up in places that would have surprised most people a few years ago.

Now, let’s get into this week’s AWS news.

Headlines
Amazon RDS for SQL Server supports Bring Your Own Media – Customers who migrate SQL Server applications from on-premises environments can now reuse their existing Microsoft SQL Server licenses, including Software Assurance, through Microsoft’s License Mobility program on Amazon RDS. BYOM is integrated with AWS License Manager for tracking license usage and compliance. Read more.

Amazon Cognito now supports multi-Region replication – You can now synchronize user and machine identity data, including credentials, user pool configurations, and federation setups, to a secondary user pool in a standby Region in near real-time. In the event of a disruption in the primary Region, signed-in users continue accessing their applications without re-authenticating, and registered users can sign in with their existing credentials. Multi-Region replication is available as an add-on for user pools in Essentials or Plus feature tiers across 16 Regions. Read more.

GPT-5.5, GPT-5.4, and Codex from OpenAI are now generally available on Amazon Bedrock – You can now use GPT-5.5 and GPT-5.4 in production workloads on Amazon Bedrock and build with Codex for AI-powered software development, with the same security, governance, and operational controls you already use across AWS. GPT-5.5 is the most capable model from OpenAI, excelling at agentic coding, data analysis, and multi-step autonomous tasks. Codex is available through the Codex App, the Codex CLI, and IDE integrations with Visual Studio Code, JetBrains, and Xcode. Pricing matches OpenAI first-party rates, and usage counts toward existing AWS commitments. Read more.

Last week’s launches
Here are some launches and updates from this past week that caught my attention:

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Upcoming AWS events
Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events as well as AWS Summits and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That’s all for this week. Check back next Monday for another Weekly Roundup!

— seb
  •  

Real-world grounding in agentic AI

The year 2026 marks a definitive shift in the AI landscape: we have moved from models that simply know to agents that do. Foundation models (FMs) — large Transformer models pretrained with massive datasets and fine-tuned for diverse downstream tasks — have moved far beyond chatbots, coding, and other digital applications. They are now used as the cognitive engines for AI agents in the physical world, where they plan, use tools, and execute multistep tasks across complex, digitally integrated environments, from warehouses and factories to transportation systems and hospitals. At Amazon, you can see the transition to this new era of "physical AI" in the debut of Project Eluna, an agentic AI model designed to transform how Amazon fulfillment centers operate. To be useful in a high-stakes physical environment, however, an agent needs to be more than fluent in natural language; it needs to be grounded in physical laws and operational constraints. In particular, we must overcome the challenge of hallucination, which, in virtual environments, takes the form of fabricated information — made-up citations, factual inaccuracies, and logical fallacies, all output with high levels of certainty. In a physical system, such hallucinations can lead to violations of reality, with detrimental consequences. For example, if an agent suggests a robotic path that ignores the momentum and mass of the items being moved, its output could be potentially dangerous to people or result in damage to products or equipment. In this article, I propose four approaches to grounding AI agents in the physical world, where "grounding" is defined as the integration of external information, including domain-specific datasets, physical principles, and numerical simulations, to contextualize a model's reasoning. All four approaches can be used separately or in combination, depending on the specific application. Practical implementation of these approaches will not only accelerate the safe and productive use of AI agents but could allow for their further expansion into new domains. Four pillars of grounding Project Eluna is an agentic AI model that lives in the cloud and assists operators who manage operations within fulfillment centers via digital dashboards. It’s designed to act with a degree of autonomy, reasoning through complex operational situations and recommending actions to operation managers. It pulls in historical and real-time data — such as the states of conveyor belts or robots — to anticipate bottlenecks and keep operations running smoothly. The four approaches to grounding AI agents that I describe here grew out of my research at the University of California, San Diego, and with the Amazon Fulfillment Technology (AFT) team, and they help ensure that agents like Eluna are physically consistent and operationally reliable. 1. Physics-guided deep learning. Traditional foundation models can learn to mimic statistical patterns in data but often fail to respect the hard constraints of the physical universe, such as the conservation of mass, energy, or momentum. In physics-guided deep learning (PGDL), we integrate first-principle physical knowledge into the foundation model in pretraining. First principles include symmetries, such as inductive biases like rotations and other transformations, and differential equations that could be used, for instance, in a robot’s motion and control. Not only does this ensure that predictions obey governing physical laws, but grounding a model in physics allows it to learn from significantly smaller datasets. If the model already "knows" the fundamental principles of dynamics, it requires less data to achieve satisfactory accuracy. 2. Uncertainty-aware reasoning. LLMs often exhibit overconfidence in uncertain predictions, which can lead to the assertion of misinformation with high certainty. For an AI agent to be trustworthy in a mission-critical setting, it must know when it does not know. Using our framework (UQ4CT), we produce calibrated uncertainty over the space of functions that map input prompts to outputs. The framework uses an approach called mixture of experts, in which the model is divided into smaller “subnetworks”, each with specific expertise. Our UQ4CT framework allows the model to dynamically align its confidence estimates with predictive correctness. Practically speaking, an agent grounded using calibrated uncertainty can halt or request human intervention when its internal uncertainty exceeds a safety threshold, ensuring reliability even when a model has been fine-tuned with relatively small datasets such as epidemiological forecasts or rare weather events. UQ4CT preserves high accuracy across five benchmarks while demonstrating over 25% reduction in expected calibration error (ECE), a measure of how well a model's estimated "probabilities" match the true, observed probabilities. Even under distribution shift, UQ4CT maintains superior ECE performance with high accuracy, showcasing improved generalizability. 3. Bridging the text-to-numerical gap. While foundation models are masters of natural language, the laws of the physical world are written in the language of mathematics and high-dimensional data, the kind used in fields like robotics, supply chain management, and finance. A trustworthy agent must translate human intent, expressed through language, into precise numerical execution without losing accuracy. Our group developed the adapting-while-learning (AWL) framework, which relies on two key mechanisms. The first is called world-knowledge distillation, where AI agents interact with simulators of the physical world to gather a range of information about what’s physically possible. This knowledge is internalized through supervised fine tuning, effectively grounding the agents’ future outputs. The second mechanism is dynamic tool adaptation, in which a foundation model calls a specialized numerical simulator when it recognizes that its original training is insufficient for the complexity of the current task. This approach is particularly useful in climate science or epidemiology. For instance, if scientists need to plan for vaccine distribution, their original model would call on outside datasets representing disease dissemination. Compared to original models without AWL, those post-trained with AWL achieved 29 percent higher answer accuracy and 12 percent better usage of simulator tools, even surpassing state-of-the-art models including GPT4o and Claude-3.5 on physical-science datasets. 4. Verifier-augmented grounding. Verifiers are software external to LLMs that can be used to ensure that the models work within the bounds of logic and reality. Our weather AI agent, Zephyrus, uses verifiers to refine the reasoning of foundation models in weather science. Zephyrus works in a “reflective” interactive loop, where the agent writes code to query outside weather datasets, observes physical results, and revises its reasoning if the output is flagged by a verifier as scientifically implausible. Another verifier, Hilbert, is used specifically for mathematical reasoning. LLMs, in general, can already generate mathematical proofs, but they need humans to verify whether these proofs are correct. However, there exist so-called proving systems, such as Lean 4, that can offer automatic verification. This has prompted efforts to build specialized prover LLMs that can generate proofs in formal mathematical language. So far, however, these provers solve substantially fewer problems than general-purpose LLMs operating in natural language. Hilbert bridges this gap by breaking complex mathematical problems into subgoals and using feedback from a separate formal verifier to validate them recursively. This process ensures that the agent’s outputs are provably correct. We’ve shown an impressive 422 percent performance improvement over the best publicly available prover LLM. Looking ahead We believe these four pillars lay a solid foundation for grounding LLMs in reality. Meanwhile, several research directions stand to deepen the connection between AI agents and the physical world. First, foundation models can be fine-tuned to interact with more complex, multifidelity numerical simulations, moving beyond function calls to agentic tools and toward an internalized sense for when and at what fidelity to invoke a simulator during reasoning. Second, uncertainty can serve not only as a hallucination detector but also as an intrinsic reward signal, training agents to explore areas of the environment where they have low confidence, high surprise, or incomplete knowledge. Third, physical laws and domain constraints can be embedded as formal verifiers during process planning. They can check every proposed action against conservation principles, kinematic limits, and safety envelopes before execution. As these techniques mature, they will increasingly work in concert: an agent that couples physics-guided learning with calibrated uncertainty and formal verification will be far more robust than one relying on any single pillar alone. Ultimately, as AI agents expand into increasingly complex physical domains, faithful reasoning and effective grounding will be the guiding principles to ensure that agentic AI operates safely, reliably, and at scale across the physical world.
  •  

Train Models Faster with JAX and MaxText Using NVFP4 on NVIDIA Blackwell

Decorative image.Pre-training frontier LLMs comes down to throughput. When training spans trillions of tokens across thousands of accelerators, every percentage point of step...Decorative image.

Pre-training frontier LLMs comes down to throughput. When training spans trillions of tokens across thousands of accelerators, every percentage point of step time can add up to days of training and substantial compute costs. Numerical precision is one of the highest-leverage knobs available, but low- bit mixed-precision pretraining is hard to get right. To address this…

Source

  •  

Bridging intent and execution in agentic systems

AI agent performance is not just a modeling problem; it is fundamentally a systems problem. A modern agent combines an LLM with a harness, software that mediates the LLM’s interaction with tools and manages the cycle of reasoning and feedback: you can think of the harness as the operating system around the model. As models improve, the performance bottleneck shifts from the model’s ability to reason to the harness’s ability to translate model intent into actions and reflect execution outcomes back to the model. In a paper we just published on arXiv, "Dissecting model behavior through agent trajectories", we formalize this bottleneck as the intent-execution gap: the mismatch between what the model intends and what the harness executes, and vice versa. For example, in trying to revise code, a model may intend to edit a single instance of a function, while the harness accidentally modifies multiple instances. We show that minimizing this bidirectional gap — without any task-specific tuning — is sufficient to achieve state-of-the-art performance across diverse agentic benchmarks, including datasets that test real-world repository patching (SWE-Pro, SWE-Verified) and interactive terminal environments (Terminal-Bench2). While the most visible components of the harness — such as the execution graph, which controls iterations over the thought-action-observation process, and tools — are natural candidates for improvement, we highlight that seemingly trivial implementation details lead to nontrivial fluctuations in performance. Factors such as environment interaction timeouts, infrastructure stability, and resource constraints also materially affect performance. Thus, benchmaxing, or reporting higher numbers on benchmarks, may not necessarily quantify underlying model/harness capability, as it is additionally influenced by the basic infrastructure parameters used during evaluations. We also introduce Simple Strands Agent (SSA), a lightweight and customizable single-agent harness designed to close the gap between the performance reported in agent documentation and the performance seen in open-source implementations. SSA achieves consistent gains in performance across multiple models and benchmarks. Finally, we show that effective agent design is not entirely model agnostic. While many principles generalize, model families differ in tool use preferences, feedback interpretation, and context sensitivity, making model-harness codesign a critical factor in achieving optimal performance. Motivations It is well established that problem-specific customizations such as tuned prompts, tailored tools, and specialized execution graphs can improve AI models’ performance in a controlled setting (fixing all other factors, such as evaluation infrastructure). However, we observed that many such optimizations fail to transfer between models. Improvements that work for one model or version often degrade, disappear, or even regress with newer models. This lack of transferability exposes a deeper issue: many optimizations implicitly overfit the behavior of a specific model. As models improve, these behaviors change, making such gains brittle and noncompounding. In the context of agents, this suggests a shift in focus: rather than optimizing for current model behavior, we should identify invariant components — design principles that remain effective across model upgrades, benchmarks, and environments. To identify such invariants, we focus on the model-harness interface — the boundary where model outputs are interpreted and executed and where execution outcomes are communicated back to the model. This interface is the primary locus of failure when agent performance degrades across settings. From this perspective, two fundamental questions emerge: Does the harness understand what the model intends to do? Is the model clear about how the harness interpreted its actions? These questions define the core alignment problem between model and harness and characterize the failure modes we analyze in the following sections. Tool-interface failures We consider the case in which the agent’s goal is code generation. Our agent primarily uses a bash tool, which provides access to the computer terminal (for example, to execute code), and a file editor to revise code. The bash tool is extremely powerful and can consume all the atomic operations of reading, searching, and editing. We make a simple enhancement to manage its outputs when they get too long. Naïvely truncating the output does not work well because the end of a command execution confirmation carries useful information such as job status and command success/failure. Instead, we contain the response length by condensing content in the middle and keeping only a limited number of lines at the beginning and the end. For reasons of efficiency and better corner-case handling in editing, we use file-editing tools in addition to bash. Our file editor is based on a string-replace mechanism that replaces existing file content with new (model-provided) content to produce edits. While string-replace works well in many cases, we repeatedly observed failure modes that expose the intent-execution gap: the model may have a clear intention, but the harness may not have enough information to execute that intention safely. In these cases, a naïve editor does not merely underperform; it can actively damage the working state by applying the wrong edit with high confidence. The first failure mode arises when the context of the model’s proposed edit appears at multiple locations in the codebase. From the model’s perspective, the requested edit may be unambiguous, because it is reasoning about a specific function, block, or error location. But if the harness receives only a raw “replace old text with new text” request, and the old text occurs several times, it cannot reliably infer which occurrence was intended. Naïvely replacing all matches is dangerous. In practice, the safer behavior is for the harness to alert the model of the ambiguity and request clarification — for example, by asking it to expand the current context such that the text to be replaced is unique. This is a small implementation detail, but it sharply improves faithfulness between intended and executed edits. A second failure mode appears when the model proposes only partial lines or short fragments for replacement. Partial-text matching is attractive because it is flexible, but it is also brittle: the same fragment may appear inside comments, string literals, neighboring expressions, or unrelated code paths. Even when the fragment is unique, replacing text that does not constitute a full logical unit — a complete line or well-bounded span — can produce malformed edits. These may be syntactically correct from the editor’s point of view but semantically unintended from the model’s point of view. We found that requiring stronger text anchors — such as exact line spans, richer surrounding context, or line-aware matching — substantially reduces these accidental edits. Put differently, the harness should not execute underspecified edit requests by guessing. Third, even when an edit is applied successfully, simply returning “edit succeeded” leaves the model underinformed about what the harness changed. This weakens the reverse side of the interaction loop: not only should the model express intent clearly, but it should also be able to verify how that intent was interpreted. To close this loop, we found it useful, after every successful edit, to supply the model with a diff file — a text file indicating what additions and deletions had been made and what text stayed the same. A diff serves as an immediate confirmation channel: the model can inspect whether the replacement landed in the correct location, whether collateral lines changed, and whether follow-up edits are needed. This seemingly minor feedback mechanism improves reliability because it converts editing from a fire-and-forget action into an observable state transition. A natural question arises: if the diff is provided after a successful edit, why do the first two failure modes require special handling? While the diff does expose unintended changes, it does so after the mistake has already been applied. At that point, the model must decide whether to roll back, repair the unintended edits, or continue execution with a potentially corrupted state. This introduces additional branching in the agent’s trajectory and forces it to spend tokens and reasoning effort correcting avoidable errors, rather than progressing toward the solution. In other words, every correction step injects additional information into the model’s context window. Note that every piece of information competes for the agent’s attention for next-action generation. Unrelated or unintended edits do not just waste tokens; they actively degrade performance by introducing spurious patterns and relationships, increasing the likelihood that the model forms incorrect associations and drifts away from the original goal. In contrast, addressing ambiguity and weak anchoring before execution ensures that edits are applied correctly in the first place. This reduces unnecessary exploration, prevents cascading errors, and keeps the context focused on task-relevant signals. In effect, the first two failure modes improve correctness at the point of action, while diff feedback improves observability after action. Both are necessary, but they operate at fundamentally different stages of the interaction loop. Reasoning A less obvious but equally important design consideration is how agents balance internal reasoning with external interactions. Chain-of-thought reasoning is clearly valuable. It allows the model to decompose a problem, plan next steps, and decide which tool to invoke. Without sufficient reasoning, tool usage becomes reactive, leading to shallow exploration, redundant calls, or poor sequencing of actions. However, excessive thinking introduces its own failure mode. When the model spends too long reasoning internally, it begins to form assumptions about the environment rather than verifying them. These assumptions may appear coherent within the model’s internal state, but they are often misaligned with the actual system state. As a result, the agent may issue poorly grounded tool calls or skip necessary validation steps altogether, creating a fundamental tension. Effective agents must continuously reconcile these two demands, and we refer to this balance as tool calling with a reasoning nudge. The idea is to encourage the model to perform just enough reasoning to decide the next action and then prioritize evidence-gathering interactions with the environment over further reasoning. Rather than extending internal chains of thought, the agent is nudged toward validating its hypotheses through tool outputs. In practice, we did not find a single “golden prompt” that reliably balances reasoning and tool interaction across all model families. For the Claude variants, we found that introducing quantitative guidance — e.g., “make 50+ tool calls” or “ideal tool call count is 100” — helps break long reasoning chains and pushes the model toward interacting with the environment. While the exact number of target tool calls is not important, it serves as a useful north star that biases the model toward action. However, in our experiments, this strong nudge was ineffective for other families, such as Gemini and Grok, which often interpret such instructions literally and make empty tool calls in order to meet the target. Such behavior reduces agent quality. Here, we find that using a flexible nudge like “You should use tools as much as possible” works just fine. The principle remains the same: we need to nudge the model to proactively use tools along with right amount of reasoning. Tool use preferences Across agents, tools function in exactly the same way, but models tend to exhibit distinct preferences in how they invoke them. For example, GPT models prefer to update code by using an apply_patch command to splice in text from a separate file, formatted in a particular way; denying them their formatting preferences hurts performance. Similarly, for Grok-4.20, a single monolithic tool for editing and viewing creates confusion, which leads to incorrect tool calls. Splitting functionality into atomic operations yields better results — even when the functionality remains unchanged. Additionally, viewing line numbers in a file helps most models, but Grok’s tokenizer and attention mechanism appeared less robust at separating prefixes from line numbers, and disabling this feature helps the view tool. These preferences are a by-product of training. This reinforces a broader design principle: agent performance is a function of not only what tools are available but how naturally those tools align with the model’s learned behaviors. A well-designed harness meets the model where it is, adapting interfaces, feedback, and interaction patterns to its strengths while still enforcing the invariants needed for reliable execution. Benchmarking study SSA is a simple harness that implements many of the principles we describe above. We evaluated it on three agentic benchmarks — SWE-Bench-Verified (n = 500), SWE-Bench-Pro (public set, n = 731) and Terminal-Bench-2 (n = 89). Each example in SWE-Bench-Verified and SWE-Bench-Pro is an open-source code repository and an “issue” to be fixed by making a code change. Terminal-Bench-2 tackles a range of programming tasks (software engineering, machine learning, security, etc.) but is not tied to a code repository. All three benchmarks have individual, static, prewritten tests for evaluating generated code. In SWE-Bench-Verified and SWE-Bench-Pro, the runs and evaluations occur in separate container images, meaning changes must be transferred into a different evaluation environment; in Terminal-Bench-2, the evaluation happens in the same container. Therefore, in SWE problems, it may be necessary to exclude irrelevant artifacts to not overly bloat the diff patch. Additionally, Terminal-Bench-2 imposes computational and agent-runtime limits that the SWE benchmarks do not. We evaluate our SSA agents using metrics standard in the field. Note that the mini-swe-agent results reported above in the SWE-Bench-Verified graph and the Terminus results reported in the Terminal-Bench-2 graph correspond to a fixed agent configuration per benchmark — the exact same prompts, tool specifications, and structural output instructions. As we discuss above, however, different model families require different reasoning nudges and exhibit distinct preferences for tool use. As a result, while SSA’s core harness remains identical, there are minimal but nonzero differences in prompts and tool specifications across model families (e.g., Claude, Gemini, GPT, Grok). Our goal in building SSA was not to optimize separate agents per model but to identify minimal, orthogonal adaptations that allow different model families to express their strongest capabilities within a shared harness framework. Terminal-Bench-2 Unlike SWE-Bench-Verified and SWE-Bench-Pro, the Terminal-Bench-2 dataset restricts the agent’s environment by limiting computational capacity (memory, storage, number of CPUs) and time (both agent and verifier run times) per project. While this is effective in limiting disproportionate use of computational resources to boost benchmark scores, it does have the unintended side effect of making the benchmark more sensitive to infrastructure choices. We observed that, given those restrictions, the following system characteristics have the most impact: Reliability of the inference backend. The inference backend’s capacity (tokens per minute and requests per minute) should be able to support all concurrently run projects for the full duration of the evaluation. High variance in invoker latency, frequent API timeouts, and retries eat into the allowed time budget, leading to more timeouts and a lower resolution rate. The number of concurrent projects run on a single node. This affects the network bandwidth available to each project. One of the first steps for an agent in Terminal-Bench-2 is to install dependencies (popular libraries like pip, torch, transformers, etc.). If the evaluation infrastructure is set up in such a way that multiple projects are run on a single node (e.g., Harbor with n_concurrent > 1), the available network bandwidth for each node is shared across all the concurrent projects. This increases the download times for dependencies, leaving the agent with less time for problem solving and a higher risk of getting interrupted before it’s done. Since the majority of tool calls involve command-line instructions, a natural way to address timeouts is to introduce a batch interface, allowing the agent to execute multiple commands in a single turn, rather than executing them sequentially. In our experiments, however, the results of this approach were mixed and correspond to one of the failure modes we describe above — the balance between reasoning and tool interaction. While batching reduces interaction overhead, it also requires the model to maintain a coherent terminal state across multiple steps, which increases reasoning complexity. For Claude models, the time taken by additional autoregressive reasoning tends to offset the gains from batching. In contrast, for other model families (such as Gemini and Grok), batch execution was beneficial, as it did not trigger additional reasoning. Overall, under constrained settings, batching commands does not consistently improve performance across all models. Given that evaluations are sensitive to such confounding factors, we next assess the upper-bound potential of the agent-model combination by relaxing time constraints. Specifically, we compare SSA’s performance on Terminal-Bench-2 under constrained settings (as shown above) and unconstrained settings, where memory and agent timeouts are removed. The unconstrained setup serves as an estimate of the achievable performance ceiling. The gap in accuracy between the constrained and unconstrained evaluations is typically 5-10%. We note that in our experiments, out of the 89 total projects in Terminal-Bench-2, a few consistently have a high timeout rate in the constrained evaluation but a high solve rate in the unconstrained setting. Those projects are make-doom-for-mips, torch-pipeline-parallelism, gpt2-codegolf, caffe-cifar-10, and train-fasttext. Experimental methodology We evaluate SSA across multiple agent benchmarks under a controlled and reproducible setup. All experiments were conducted on an AWS PCS cluster using c7.48xlarge instances, with maximum concurrency set to 10 to balance throughput and system stability. For model access, Claude models were served via Amazon Bedrock (production capacity), while OpenAI, Gemini, and Grok models were accessed through their respective commercial APIs. We enforced strict evaluation hygiene. Internet access was disabled for SWE-Bench-Verified and SWE-Bench-Pro runs, while it was enabled for Terminal-Bench 2 due to its benchmark design. For SWE-Bench-Verified and SWE-Bench-Pro, we used the standard benchmarking Docker environments, which include repository state up to the point of the current code revision. This allows agents access to the relevant history of the codebase while ensuring no access to future revisions. Evaluation-specific issues In SWE-Bench-Verified, instances such as astropy-8872 and astropy-8707 fail even with flawless code patches due to setup inconsistencies and require fixes in the evaluation environment. Additionally, some psf_requests instances can fail intermittently due to external test dependencies (e.g., nonresponsive URLs), requiring manual patching for reliable evaluation. For SWE-Bench-Pro, evaluations were executed on Amazon ECS. Due to environment-specific assumptions, a small subset of tests — 3 out of 731 instances — consistently fail when run on AWS infrastructure, resulting in an approximate 0.41% ceiling loss across all SSA evaluations. Finally, to minimize information leakage during agent runs in Terminal-Bench-2, hidden tests are introduced into the Docker environment only after the agent has completed its execution, ensuring that the agent has no direct access to them during problem solving. Note that internet access in Terminal-Bench 2 does introduce a possibility of solution leakage, but a manual review of trajectories didn’t reveal any instances of the model trying to copy solutions. Model configs To ensure reproducibility, we used public documented configurations from release/model cards wherever available. Specifically, Claude Opus 4.6 and Claude Sonnet 4.6 were used with adaptive thinking and max effort across all benchmarks (except when Sonnet 4.6 was tested on Terminal-Bench-2 with thinking disabled). Opus 4.5 used high effort and no thinking across all benchmark runs (except in Terminal-Bench-2, where Opus 4.5 has thinking enabled with 128k budget tokens). Sonnet 4.5 was used with an interleaved-thinking budget of 200k, Haiku 4.5 with a 128k budget, and Sonnet 4.0 with a 200k budget across all runs. Both Gemini 3.0 Flash and Gemini 3.1 Pro used thinking_level high and temperature 1.0 across all runs. Every GPT model used reasoning effort xhigh for all benchmarking runs. With Grok, we used the grok-4.20 reasoning variant for all runs with default configs. Detailed config files for every experiment are included in the SSA package. Conclusion We show that bridging the intent and execution gap in agent harnesses is critical to extracting state-of-the-art performance out of frontier models. Well-chosen editing tools, feedback from tool application, and management of tool-output lengths improve performance across all model families. On the other hand, models exhibit distinct preferences for different tool interfaces, and an effective harness should leverage them instead of trying to uniformly impose the same interfaces across all model families. We open-source all elements of our harness — the agent logic, tools, and prompts, as well as model configs, for easy reproducibility in the SSA package. Acknowledgments: Luke Huan and Anoop Deoras
  •  

Try the new console experience in Amazon Bedrock, optimized for Anthropic- and OpenAI-compatible APIs

Today, we’re announcing a new console experience in Amazon Bedrock for you to experiment, iterate, and scale with the latest AI models on Amazon Bedrock’s next-generation inference engine built for high performance, reliability, and security. This console has a refreshed workflow optimized for bedrock-mantle endpoint, which supports the latest GPT, Claude, and open-weight models with the OpenAI Responses API, OpenAI Chat Completions API, and the Anthropic Messages API.

The new console experience makes it simple to find the right model and move quickly from evaluation to production.

  • New model card: You can browse the full model catalog, compare them side by side on capabilities, modality support, context window, and applicable service quotas in a single view, removing the need to stitch together documentation, and limit calculators.
  • Project-based work: You can make a project to run evaluations and review usage insights in one streamlined workflow that mirrors the lifecycle of building a generative AI application.
  • Live documentation: You can use project-aware live documentation: code samples, SDK snippets, and API references are automatically prefilled with your project variables. You can copy a snippet straight from the console into your application and run it without modification.

How to get started
You can try a new experience by choosing Try the new Bedrock Console from within the Amazon Bedrock console, or by using the new console link directly.

You can find a project-based dashboard to show inference requests and error by range of recent dates, recently used models, and the project list. You can create a project, assign models, configure API keys, and start making inference requests in minutes.

A new model catalog shows the latest GPT, Claude, and open-weight models that are supported on the bedrock-mantle engine. You can see the details of features, tokens, pricing, input/output, pricing information, and Regional availability. You can also compare up to 3 models in a single view.

When you choose the project dashboard, you can see the models used in the project, the distribution of your token usage such as total token usage, token usage per minute, inference requests per minute, and tokens per inference request. This can inform your model selection, prompt optimization, and workload consistency decisions.

You can select up to 3 models to start evaluating to compare responses side by side with the same prompt.

To build your application in the project, choose Getting started. You can migrate existing code, build a new app with the Anthropic or OpenAI SDK, or connect an AI coding assistant to Bedrock.

Choose the API & SDK, your SDK (either Anthropic or OpenAI), your preferred programming language, and your authentication method. It shows your environment code to run these in your terminal for a quick test, or save to a .env file for your application. You can also send your first request with sample code snippets to verify your setup.

When you choose Clients, you can select the AI coding agent source such as Claude Code, Cline, Codex, Cursor, or OpenCode that you want to connect to the bedrock-mantle engine. It provides instructions on how to install the AI agent, use your AWS IAM credentials or use a Bedrock API key, set environment variables, and route requests from each AI agent through Bedrock.

To learn about Anthropic- and OpenAI-compatible APIs, choose Live API docs. You can choose Anthropic API Protocol for access to Claude model features like the Messages API or OpenAI API Protocol for access to features like Responses API.

For example, when you choose OpenAI Response API, it retrieves a model response with the given model ID. These API references are automatically prefilled with the project’s selected model ID, Region, bedrock-mantle endpoint URL, and API key reference, and they update in place as you change models or settings.

You can also choose the existing Bedrock console to manage fully-managed features such as Agents, Knowledge Bases, Guardrails, fine-tuning, or the InvokeModel and Converse APIs to run on the bedrock-runtime endpoint.

Now available
The new console experience is available in all AWS Regions where the bedrock-mantle endpoint is offered: US East (N. Virginia, Ohio), US West (Oregon), Asia Pacific (Jakarta, Mumbai, Sydney, Tokyo), Europe (Frankfurt, Ireland, London, Milan, Stockholm), and South America (São Paulo). Check the full list of Regions for future updates.

Give the new console experience a try in the new Amazon Bedrock console and send feedback to AWS re:Post for Amazon Bedrock or through your usual AWS Support contacts.

— Channy

Updated on July 13, 2026 – Fixed the first console image to reflect the latest UI.

  •  

Introducing the Google Colab CLI

Google has announced the Google Colab Command-Line Interface (CLI), a new tool that allows developers and AI agents to connect local terminals to remote Colab runtimes for frictionless execution. The lightweight CLI enables users to easily request high-powered GPUs, run local Python scripts remotely, and seamlessly retrieve artifact logs or models like fine-tuned Gemma 3 adapters. By integrating directly into standard terminal environments, the tool is highly programmable and ready to be used by AI agents such as Antigravity or Claude Code to manage complex machine learning pipelines.
  •  

Video Friday: Watch This Running Robot Not Fall Down Stairs



Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO

Enjoy today’s videos!

It’s been a while since a humanoid robot video actually impressed me, but the beginning of this does.

Hard to know how much of that recovery was luck, though.

[ Deep Robotics ]

When you’re very confident in your MPC-based balance controller...

I feel you, buddy. And thanks for posting this. We’ve all been there, in one way or another.

[ DARoS Lab ]

GENE01 designed from scratch and sent to batch production. Two scalable lower bodies. Physical AI deployed for motor control and world-action modeling. All in three months. Generative Bionics is running.

[ Generative Bionics ]

Alex, the newest humanoid robot built entirely by ‪IHMCRobotics‬, takes its first steps outdoors! This was a significant milestone for our team, especially because Alex is the first humanoid robot developed entirely by IHMC Robotics to venture outside the lab. These outdoor trials were conducted in preparation for a demonstration in Maryland, where Alex later successfully walked completely untethered.

[ IHMC ]

Built on the Enlight platform, Flexiv Mico is a compact dual-arm system engineered for safe, seamless collaboration in any workspace.

[ Flexiv Robotics ]

Is it weird that I’m jealous that robots can have feet that are swappable?

[ Boston Dynamics ]

Midweek at ICRA 2026. Cable-climbing robots that work as a squad: CCRobot-S is a team of robots with reconfigurable cable-driven manipulation that collaboratively inspect and maintain long-span bridge stay cables. Parallel operation for speed, morphological reconfiguration for reach.

[ IEEE Transactions on Robotics ]

I would love to know the story behind this odd choice of hat.

[ Robotis ]

How did Atlas learn football—and why? Go behind the scenes of School of Football and discover a glimpse of the Next of robotics.

What I really want to know is, what kinds of things is it possible to do in football (soccer) when your robot’s joints are not constrained by biology?

[ Boston Dynamics ]

  •  

AI for inclusive and resilient agri-food systems: Potential ways forward  

Global agri-food systems are under growing strain. Even though the world produces enough calories to feed more than the world’s entire population, one in eleven people – or nearly 700 million people – still face hunger. Climate shocks, fragile supply chains, and labour shortages increasingly threaten the resilience of agri-food systems worldwide. As pressure on farmers and supply chains intensifies, artificial intelligence (AI) is a promising tool to ensure that all stakeholders – including vulnerable populations – can benefit from the transition towards more resilient agri-food systems. 

Agri-food systems face growing strain – AI tools offer solutions 

The agricultural sector plays an essential role in economies and societies around the world, providing communities with reliable, quality food and tens of millions of jobs. Yet the global agri-food system is under pressure, as various issues, from climate challenges to workforce shortages, put strain on farmers and global supply chains.  

AI offers opportunities to address the needs of farmers and other actors in the agri-food system across diverse local contexts. And AI is already transforming agriculture and related supply chains. It helps farmers predict droughts, reduce pesticide use and identify crop diseases before they spread. It supports buyers, traders and retailers in making decisions regarding product availability, quality and prices. AI-enabled precision spraying can reduce pesticide use by up to 30% without compromising yields. Drought-tolerant traits identified using AI in sorghum and chickpea crops boost yields by up to 25% during dry seasons. And the Global AI Hybrid Rice Platform shortens breeding cycles by predicting optimal parent combinations.  

However, challenges persist in the adoption of AI and other digital technologies. Access to these promising tools is starkly uneven, and so is the distribution of information throughout supply chains. In Australia, nearly 96% of farmers use digital tools, whereas in Chile, just 12% do. In addition, issues persist in data interoperability between different stakeholders and jurisdictions, which hinder data sharing opportunities and limit the impact of digital tools. Digital tools need to be accessible, trusted and designed to respond to local user conditions.  

Cybersecurity must also be treated as a foundational prerequisite for AI-enabled agri-food systems (e.g. through secure-by-design approaches). Without robust protections from cyber threats and concrete actions such as building digital literacy and aligning policy agendas, there is a risk that the potential of using digital technologies will not be realised or will have adverse consequences. At the same time, the lack of agri-food-specific AI governance could create regulatory uncertainty, making the case for targeted frameworks that promote the trustworthy deployment of AI while accounting for the sector’s unique characteristics. 

These issues were at the centre of a flagship session on AI for Inclusive and Resilient Food Systems, co-hosted by the Kingdom of the Netherlands and the OECD at the India AI Impact Summit in New Delhi in February 2026. The session leveraged the work of the Summit’s Working Group on Economic Growth and Social Good, co-chaired by India, Indonesia and the Netherlands. Examples from these three countries and field‑level research highlighted where AI is already showing great promise and where critical bottlenecks remain.  

These insights point to three areas where deeper international co-operation and policy analysis facilitated by organisations such as the OECD could help countries make meaningful progress. Specifically, the OECD’s Global Partnership on AI (GPAI), could take next steps and examine where AI is delivering results in agri-food systems across various regions and what it will take to ensure benefits are widely shared. 

The opportunity: How AI can help build efficient and resilient supply chains for global food security  

Efficient and resilient supply chains are cornerstones of the agriculture sector and global food security. Strengthening global food security is a strategic priority for countries like the Netherlands because reliable, sustainable and affordable food systems are essential for societal stability and economic development, particularly in vulnerable regions. After the United States, the Netherlands is the world’s second-largest exporter of agricultural products. The Dutch have already seen tangible results from AI applications in agriculture. According to research from Wageningen University, AI tools for advanced greenhouses have yielded water savings of up to 90% through smart irrigation, compared to traditional open-field systems. 

The challenge is that successful solutions in countries like the Netherlands cannot always be exported to other countries due to the diversity of agricultural contexts worldwide. In view of this complexity, partnerships are an important tool. Examples from the Netherlands highlight the importance of co-creation as a vital strategy for tangible results. The Netherlands works with other countries to develop locally relevant AI solutions that are inclusive and accessible to farmers through knowledge sharing, capacity-building and co-creation. 

The use of AI can also help reshape how companies engage with stakeholders across supply chains and help to manage risks. Tools intended to improve efficiency (e.g. automated advisory systems, remote monitoring, or worker feedback platforms) can enhance visibility and responsiveness, but they can also displace meaningful human engagement if not carefully implemented. Managing these trade-offs is therefore essential to ensure that AI contributes positively to sustainability efforts. The OECD Due Diligence Guidance for Responsible Business Conduct and AI-specific due diligence guidance provide a practical framework for managing these risks. It promotes a risk-based, continuous approach grounded in stakeholder engagement, with step-by-step operational guidance to identify, prevent, mitigate, and account for adverse impacts. 

Despite promising tools and guidance, in practice, there are persistent problems with AI adoption across agri-food systems and supply chains, including among farmers. Below are three key challenges which would benefit further from OECD and GPAI analysis and co-operation. 

Solving the data problem: improving interoperability and shared agricultural data  

Across regions, the most consistent barrier to effectively using AI in agriculture is data: not enough of it, not widely shared, not of high enough quality and often not interoperable. Addressing these challenges is at the core of many initiatives of the Dutch Ministry of Agriculture (LVVN), as well as EU initiatives such as the Common Agriculture European Data Space and the European Digital Infrastructure Consortium for Agri-Food.  

When agricultural data remains siloed between ministries, supply‑chain actors or countries, AI tools underperform in real‑world conditions. This challenge surfaces in examples such as global crop mapping efforts, which struggle when key national data is unavailable, and contrasts sharply with locally tailored tools. Data interoperability is critical not only for enhancing transparency and traceability but also for enabling benchmarking that drives competitiveness and for unlocking new market opportunities. 

In the cocoa industry, for example, which is suffering heavily from climate change, researchers from Wageningen University built a chatbot in the farmers’ local languages to identify plant diseases using computer vision trained on local data in the form of images. It worked because it was built with locally sourced data using the farmers’ perspective, not the researchers’. This shows that not only are data availability, quality, and interoperability important, but trust also plays a critical role in AI uptake. The same is true for AI literacy. For example, an AI tool’s success in the field depends greatly on a farmer’s ability to generate, manage and apply high‑quality, context‑specific data. This could include skills spanning data literacy (e.g., collection, quality control, labelling), applied digital skills (e.g., the use of sensors), and general skills for successfully interpreting and assessing AI outputs.   

Building shared and trustworthy data infrastructures is a challenge well suited to multilateral co-operation. The OECD and GPAI could help countries examine which types of agricultural data are most critical for AI uptake, how to enable cross-border interoperability, and how governance, incentives, and standards can encourage responsible data sharing without disadvantaging farmers or exposing sensitive information. Such analysis would complement ongoing OECD and GPAI work on AI governance and help countries design data ecosystems – for example, through bilateral agreements – that support resilience rather than reinforce fragmentation. 

Small but mighty: ensuring AI works for small farmers through local relevance and co-design  

AI will only advance global food security if it also works for small farmers – who grow roughly one‑third of the world’s food but are also most vulnerable to climate and market shocks. Many tools fail to provide the expected socio-economic value because they are built for ideal conditions or assume a baseline of digital maturity that simply does not exist. The most effective way to address these issues is to begin with the farmer’s perspective. 

Examples from India illustrate this opportunity. A voice-first AI tool developed by the Government of India, called BharatVistaar, provides farmers with a breadth of agricultural advisory information on subjects ranging from shrimp cultivation to pest control. It comes as a simple phone call or text message from a chatbot, in their local language, with no smartphone required. This accessible, low-tech solution shows how AI can benefit farmers who lack access to complex technology or reliable internet. 

This is an area where the OECD and GPAI’s evidence base and global reach could offer unique value. Through comparative analysis and its global network of experts, the OECD could help countries better understand what farmer-centred AI design looks like in practice; which models scale across different agricultural contexts; and how to build trust by aligning tools with real-world decision‑making. Sharing case studies through the OECD.AI Policy Observatory, sharing concrete field-level agri-food system AI tools through the OECD.AI Catalogue of Tools and Metrics, contributing policies to the OECD.AI Policy Navigator, leveraging the OECD.AI Policy Toolkit, and sharing local lessons at GPAI meetings could help countries learn from each other’s successes and pitfalls. 

Moving from pilots to scale 

Scaling AI solutions in agri-food settings is not easy. Solutions that perform well technically often falter in field conditions, and many projects remain stuck as promising pilots that never achieve systemic impact. Factors such as infrastructure gaps, cybersecurity threats, institutional capacity, financing constraints and lack of local adaptation all limit the path from prototype to widespread adoption. 

Scaling and resilience in agri-food systems must extend beyond agriculture itself to the AI infrastructures that underpin it, as cyberattacks can render systems unavailable – particularly impacting smallholder farmers. An emphasis on accessibility without security could lead to unsustainable outcomes, making cybersecurity a critical precondition for success. 

Countries like the Netherlands are working with partners to co-create AI solutions that are secure and deeply adapted to local ecosystems to be successful in scaling up. Indonesia also offers a compelling example. With more than 17 000 islands with varied soil conditions and uneven infrastructure, the country sees AI as essential to developing resilient agriculture and has integrated AI into its national strategy for climate-resilient agriculture to combat scaling challenges related to its diverse geography.  

A structured examination of what it takes to scale AI responsibly for agriculture – across geographies, farm sizes and value chains – could provide actionable insights for governments worldwide. This could include analysing enabling policies, public‑private partnerships, field-level capacity building, and pathways for adapting successful models to regions with differing levels of infrastructure development. 

With its cross-country reach and evidence-based approach, the OECD, through GPAI, is well-positioned to convene comparative analysis of scaling and cybersecurity challenges, identifying practical levers to help countries move from fragmented experimentation to system-wide adoption. 

Looking ahead: tackling these challenges through the Global Partnership on AI  

AI has the potential to transform global agri-food systems, but technology alone cannot deliver this outcome. The choices made around governance, access and partnerships will determine whether AI strengthens resilience broadly or deepens existing divides. Advancing AI for sustainable agri-food systems depends on addressing the three imperatives discussed above – ensuring access to high-quality data by building interoperable, trusted data ecosystems; designing farmer-led, locally relevant solutions; and creating the conditions to scale securely and sustainably through robust infrastructure, governance, and strong cybersecurity to safeguard system resilience. These topics are at the core of the Dutch Ministry of Agriculture’s (LVVN) priorities and must be complemented by fit-for-purpose policy frameworks alongside viable financial models, knowledge exchanges, and innovation partnerships to enable effective and inclusive adoption. 

The OECD works with countries – including the Netherlands – at various levels of economic development to establish AI governance based on the OECD AI Principles. The OECD.AI Policy Navigator gives access to AI policy initiatives in the agriculture sector, covering more than 2,000 AI policies and initiatives across 80 jurisdictions. Any country can use it to benchmark and strengthen its approach. Further analysis of the challenges and opportunities for AI in agri-food systems could be undertaken as part of initiatives such as GPAI and groups like the OECD-FAO Advisory Group on Responsible Agricultural Supply Chains.  

Interested countries and partners are invited to contact ai@oecd.org to explore this topic further.  

The post AI for inclusive and resilient agri-food systems: Potential ways forward   appeared first on OECD.AI.

  •  

NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents

Illustration showing Nemtron 3 Ultra.Single-turn chatbots are evolving into long-running agents that can reason, maintain context, use tools, and run efficiently across many turns to complete...Illustration showing Nemtron 3 Ultra.

Single-turn chatbots are evolving into long-running agents that can reason, maintain context, use tools, and run efficiently across many turns to complete complex workflows. However, these multi-agent workflows cause token counts to grow quickly. Agents plan, call tools, invoke sub-agents, receive information, and then pass history, outputs, and reasoning steps back into the model…

Source

  •  
❌