The past few weeks have probably made any company running a flock of agents stand up and take note. We now have models that don’t just reward-hack but (as part of an agentic system with sufficient computational resourcing) can do so at astonishing scale and capability.1 Since the initial reports, evidence of misaligned behaviors by agents driven by frontier lab models has continued to accrue, including additional cybersecurity incidents and frank attempts to hack third-party systems.2 Those same models, or in some cases purportedly tamer versions, have made impressive, if sometimes variable gains on benchmarks, including those that aim to test general intelligence.3 They can also increasingly do many more things on a computer without specialized software connectors (like a human user might). Depending on the day, the frontier labs use these developments to argue that we’re headed toward a bright future of unprecedented productivity, mass unemployment of the white-collar labor force, or Skynet… and sometimes all three at the same time.

Faced with these competing futures, public debate has increasingly focused on whether there is a need for the frontier labs to improve their oversight of product quality and safety, for embedded third-party evaluators to assess security breaches by said products, or for increased regulation of the industry. Our answers are: without question; great idea, but ideally not by captive channel partners that help sell the same products they evaluate; and quite possibly, with the devil being in the details.

What the hyperbole crowds out

The problem with all this hyperbole is that it elevates a very focal set of problems, for a relatively clear set of actors, with actionable solutions, even as it distracts from the real challenges for most companies looking to do something productive with AI. At the same time, with notable exceptions (cybersecurity being one), there is surprisingly little evidence that the bleeding-edge models causing all the fervor are even necessary to do many of the kinds of work that most companies and users want or need.4 We don’t dispute that newer models open up new capabilities in certain instances, but even when they do perform better than a simpler model, evidence that any performance gains are worth the substantially greater cost is slimmer still. Separately, in at least some cases and possibly many, the relatively traditional software that houses the model may be just as important or more important for performance than the model ‘brain’ inside.5

What actually matters

Regardless of the direction of public debate, we’re decidedly AI optimists. We’ve built our own business on AI, implemented it at the companies we’ve operated, and worked with clients on AI product and implementation strategy. Through that work, we find that the most important considerations for leaders looking to implement AI are far more prosaic than the headlines would suggest, being fundamentally practical, real, and consistent:

For companies building agentic or AI tools, the considerations are similar but redirected to the level of product strategy and economics. Of course, in pursuing solutions for these considerations, neither AI builders nor companies implementing AI systems can afford to ignore the risks of misaligned behaviors, particularly now that they have been observed at scale. Nonetheless, the solutions are far more likely to sit in the more tedious work of robustly engineering, evaluating, and observing AI systems before and after they are pushed to production, rather than in the next frontier lab model release.

Models aren’t friends, they’re infrastructure

For both AI implementers and AI builders, which models (and model providers) to use should primarily be a question of whether value is commensurate with cost and risk. In our own experience, even small models (including some we run on our own hardware) can perform very well — even as part of workflows that require meaningful judgment and reasoning — particularly when optimized through instruction and fine-tuning.6

Whether implementing or building, one practical part of the solve is to treat models as infrastructure: pick the least costly model that does the work and avoid building to any single provider’s distinctive features. The corollary is that organizations should be ready to swap models or providers when the economics change or the risk profile can be improved upon. Available and emerging tools make this increasingly feasible7, and while making such a swap is never going to be as simple as flipping a switch, the capability to do so is a realistic goal for most companies and AI products.

Managing risk requires systems-level thinking

Since the first LLM-powered chatbots appeared on the scene, the distinction between what they can do and what they can do reliably has been a topic of discussion.8 Recent events require an addendum: it’s not just what they can do reliably but also what they can do safely.2 Far more important than whether the latest model is powering an agentic system is whether that agentic system is integrated into an organization’s operations and systems in a way that can be trusted at scale.

Most general-purpose coding or knowledge-work tools from the frontier labs can be building blocks inside such trustable systems, but on their own do not provide the kind of surety that most businesses require. For most AI implementers (and AI builders by definition), this means that trustable systems may need to be composed rather than simply purchased off the shelf. While the considerations will vary by use case and organization, general principles apply: agentic systems should work in the least permissive environment possible to achieve the goal. This is nothing new as it is essentially the ‘Principle of Least Privilege’9 that usually governs human users’ access, with the notable caveat being that many agentic systems can be more capable than human users at pushing the limits of their computing environment.

Evaluation, observability and oversight as non-negotiables

Evaluation against an organization’s own work is what makes the cost and risk of an AI system, and therefore its value, predictable. Observability (tracing actions, tool calls and token spend for each task) keeps them predictable in production. Oversight, scaled to the consequences of an action, keeps a person accountable for decisions that can’t be undone. None of these depends on which model sits inside the system, but blanket use of more capable models probably makes them more critical.

The frontier labs will continue to announce models that they claim are closer to AGI, more dangerous, or both. For a company trying to use AI effectively and responsibly, such announcements are generally a distraction from what actually matters on the ground: what the work requires, what value an agentic system returns, and how the company will know when something goes wrong.


Josh is co-founder of Kynetyk, where he writes about AI, builds products at the intersection of AI and human experience, and helps companies design AI strategies that actually scale. Reach out at josh@kynetyk.ai.

  1. Misalignment and reward hacking are present with all models, and neither is new. What changed is the combination: a highly capable model able to stay coherently intentional across a long task, wrapped in a harness, given substantial computational resources, pointed at poorly secured tools, and overseen loosely or not at all. In July and August 2026, OpenAI, Anthropic and Meta each disclosed that a model had reached a real external organization’s production systems from inside an evaluation environment it believed was isolated; in one Anthropic case the model recognized it had reached production and continued the intrusion anyway. See the Cloud Security Alliance’s research note for the three disclosures together, and METR’s work on task-completion time horizons for the long-horizon coherence that makes this class of incident possible. ↩

  2. Anthropic’s alignment assessment of recent cybersecurity incidents (9 September 2026) reports a fourth case of a Claude model gaining unauthorized access to a third party’s systems during an evaluation. It identifies two behaviors that recur across all four cases: biased reasoning, in which the model dismissed evidence that it was on the real internet, and recklessness in narrow pursuit of its task. OpenAI’s misalignment reporting framework (16 September) discloses six more incidents, and OpenAI has since notified “dozens” of third parties of improper agent activity while it works to establish the full scope (Reuters). Cases reported since include agents that gained unauthorized access to four Australian government websites, one of them a Medicare statistics portal (The Guardian), and agents that hit a UN Trade and Development data site more than 16,000 times, brute-forcing API fields to get around its restrictions (The Verge). In the UN case, attribution rests on a researcher’s analysis rather than confirmation from OpenAI. ↩ ↩2

  3. See the ARC Prize leaderboard for ARC-AGI-2 and ARC-AGI-3, and OSWorld 2.0 for computer use. ↩

  4. The gains are uneven across benchmarks, and the routine end of the range is largely finished. GPT-6 Astra reaches 99.9% on ARC-AGI-3, against a human baseline of roughly 48%, effectively saturating it. Set that against Artificial Analysis’s evaluation of the same model, which has it regressing on agentic coding — 68% on DeepSWE where its own predecessor scored 72% — and shedding roughly 45 Elo on GDPval-AA v2, a benchmark built around economically valuable work. Analytical quality rose on AA-Briefcase while presentation quality fell, the older model still leading there. Further down the range the picture is flatter still: on the Berkeley Function Calling Leaderboard, simple single-call tool use is essentially saturated across frontier models, and the separation only appears in multi-turn. Most business work sits at the saturated end. ↩

  5. Two examples. ARC Prize scores GPT-6 Astra at 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a provider-adapter harness that preserves reasoning state between requests and compacts longer conversations. Although the near-perfect run was the cheaper of the two, at $18,817 against $26,098, it’s unclear how smaller models would perform in the same provider-adapter harness. Anthropic’s April 2026 postmortem on Claude Code quality is the same lesson from the other direction: three product-configuration changes — default reasoning effort lowered from high to medium, a caching bug that cleared the model’s thinking every turn rather than once, and a system-prompt instruction capping response length — degraded output for weeks. The weights never changed there either. ↩

  6. Beyond our own experience: NVIDIA’s Small Language Models are the Future of Agentic AI argues that SLMs are “sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems.” Typesafe AI’s Jev pushes the same argument further, building for machine-native integration rather than conversation on the premise that “the bottleneck isn’t raw intelligence.” ↩

  7. OpenRouter brokers one API across providers with fallbacks and automatic routing; AWS Bedrock serves many model families inside existing account and network boundaries. Either turns a swap into a configuration change rather than a rebuild. Open-weight models remove the provider question altogether, and open-source harnesses and workspaces now replicate much of the coding, chat and knowledge-work tooling the providers wrap around their own models. ↩

  8. τ-bench measures this with pass^k, the rate at which an agent succeeds on the same task across k attempts. Success falls sharply as k rises, even for models that complete the task once. ↩

  9. NIST defines least privilege as granting “each entity … the minimum system resources and authorizations that the entity needs to perform its function” (SP 800-53 Rev. 5). OWASP applies the principle to LLM agents under Excessive Agency: limit the tools an agent can call, the functions those tools expose, and the permissions they hold, and require approval for high-impact actions. The Anthropic incidents in footnote 2 started with a misconfigured evaluation environment that connected the model to the open internet. ↩