Tag: bedrock
All the articles with the tag "bedrock".
-
Opus 5.5 Leads a Four-Model OWASP Java Test by 7 to 9 Points. Kimi K3, GLM 5.3, and GPT-6 Sol Tie.
A 660-case OWASP Benchmark Java test at maximum reasoning effort across four models on Amazon Bedrock, with accuracy, missed flaws, false alarms, token use, and cost per correct answer.
-
AgentCore Code Interpreter: When the Agent Has to Actually Run the Math
Updated:AgentCore Code Interpreter ran a Pareto frontier over 528 models and matched the known-good result; tampering collapsed 13 results to one.
-
AgentCore Memory: What an Agent Remembers When the Session Is Gone
Updated:AgentCore Memory recalled preferences across microVM sessions and isolated three tenants, but asynchronous extraction can briefly return stale values.
-
AgentCore Runtime: Where an Agent Actually Runs
Updated:AgentCore Runtime hosts each agent session in a managed microVM behind an HTTP server contract. Seven test sessions ran with zero errors.
-
Browser Use vs. AgentCore Browser: Two Managed Browsers, and Which Layer Each One Wins
Updated:AgentCore Browser supplies governed Chromium; Browser Use adds a natural-language agent and cached reruns. Eleven experiments expose the tradeoffs.
-
Multi-Agent Memory: Agents That Actually Share Context
Updated:Two agents shared context through one AgentCore Memory store while actor namespaces isolated a third. Semantic extraction required a two-turn exchange.
-
AgentCore Observability: What Traces Knew That Logs Didn't
Updated:Eight AgentCore debugging lessons cover dead runtimes, missing spans, cross-agent loops, protocol noise, cold starts, and off-runtime telemetry.
-
What You Build on the AgentCore Harness
Updated:Bedrock models reason; AgentCore hosts, remembers, authorizes, and observes. Four use-case architectures show how the managed harness fits together.