Why the “SaaS Is Dead” Vibe is Dead Wrong…at Least in Biopharma

Through a combination of increased exposure to and improved capability of general purpose AI tools, 2026 has probably been the year where appreciation of the potential of AI tools went mainstream (along with appreciation of other less exciting things like token costs). For a while, the general sentiment was that the SaaS business model was toast - if a coding agent couldn’t do a task, it would just write code or vibe code a system of record on the fly that could.1 While that sentiment has started to fade and was probably never a great idea to begin with, AI SaaS builders’ response to the potential of AI is still immature, particularly in cases involving regulated industries and complex or highly governed workflows as in biopharma.

Where AI Tooling for Biopharma Stands Today

We’ve recently surveyed AI tooling for clinical development workflows, an area of particular relevance for our biopharma clients and have observed that immaturity firsthand. If one wants to integrate AI into highly regulated and governed clinical development workflows the options in many areas are still AI tools scoped to the use of a legacy system or platform or general purpose tools and platforms that leave agent and workflow definition, evaluation, and (where required) validation to the user or internal IT functions.2 While the former is useful, meaningful units of work often involve coordination across multiple systems, and for the latter, many, and likely most biopharma companies are not equipped for that kind of ‘building’ - at least when the system in question needs to be used at scale and does not intrinsically drive enterprise value creation.

To be clear, we don’t argue that general-purpose AI tools (eg, coding and cowork/copilot-style agents) aren’t important; in fact, we use them every day. In a small business or start-up they can do a lot of critical work when used by contributors who ‘drive’ them well and are simultaneously accountable for business outcomes and business execution. The problem comes with scale: when companies need robust systems, overseen by many individuals with varying levels of ‘AI fluency’, and where there is distributed accountability for process and outcomes. In these cases there needs to be a better way, and we think that means there is significant opportunity for AI SaaS builders who get it right.3

How SaaS Builders Can Embrace AI for Growth

It almost goes without saying that coding agents have further commoditized code and the ability to produce it. But the value of a software product hasn’t rested solely in its code for a long time. What buyers are really paying for is the ability to outsource testing, maintenance, security hardening, and to some extent implementation to a trusted third party. Companies could already self-host any number of open-source platforms, and most don’t, because doing all of the work that comes along for the ride simply doesn’t make good business sense. If anything, the cybersecurity concerns raised by frontier model capabilities make that line of reasoning even clearer.4 However, relying on those reasons alone isn’t enough; as more work is done by agents, software should evolve with that reality; companies and products that do will find opportunity and those that don’t will increasingly be left behind. In this respect, we see at least two distinct near-term opportunities for SaaS builders: reimagining systems of record, and building a new class of ‘agentic services’ scoped to a meaningful unit of work.

Systems of Record at the Intersection of Human-AI Work

For those who develop systems of record (eg, accounting systems, CRMs, eQMS, RIM platforms), we think the opportunity isn’t about adding a tightly scoped AI agent as a feature (although many appear to be doing this) but rather making that system the place where meaningful AI and human work intersect. As business productivity increasingly involves both AI and human contributors5, a system of record can and should be the single source of truth for both. It can also be a ready-made solution for effective human oversight of agentic systems: most already include the well-developed GUIs, dashboards, logs and audit trails that humans need to effectively review, and actually be accountable for, AI work product.

Where many of these systems need to evolve is in how agents integrate with them. The integrations to do so exist in some cases, but connecting them today is risky and can compromise the system as the source of truth for humans and agents alike.6 Systems of record should instead treat the ingestion and staging of untrusted or semi-trusted AI proposals as a first-order concern, and expose well-documented API endpoints that allow agents to make such proposals as part of networked systems, with traceable human review before critical changes are accepted.

Agentic Services as a New Product Category

The second opportunity is a product scoped not to a system but to a unit of work that a buyer would recognize. As noted above, most of the AI tooling we surveyed improves the use of a single system without addressing the real job to be done, which often spans multiple systems and information sources. The popular vision for this kind of product is an AI coworker that takes on a job the way a person would. We think that vision is hard to deliver, even with models that have been described as approaching AGI, at least in the industries we tend to work with. Long-running agents often lose the thread of the task, try to do too much with too little consistency, and drown in their own context.7 The work also demands accountability, which AI cannot deliver, and judgement, where we find even the most capable models frequently fail when on autopilot.

The scope that works is smaller than a job but significantly larger than a single task: a unit of work narrow enough that a well-engineered system can execute it reliably, and substantial enough that handing it off is worth paying for. The work that companies already outsource to third-party service providers, or could outsource, is a natural place to start. The division of labor between oversight and execution is already familiar to the people involved, and outsourcing of that work already represents meaningful cost centers that can be optimized across a P&L. The product should also be built like a service provider: multiple agents that each perform, verify, and deliver a bounded piece of the work as one integrated system, with the buyer overseeing the output as they would a vendor’s. We call these products agentic services and avoid “agents” because that word suggests the single coworker. A coding agent with a well-written skill can do some of this work, but that approach is generally hard to scale, hard to govern, and, in our experience, unreliable over time. Moreover, results depend on who is driving, and driving takes constant manual steering, which can reduce a highly trained professional to a console operator whose job is keeping an agent on course.8

AI SaaS builders developing agentic services need to supply what that approach lacks, which means defining the work (granular workflows, configurable scope, the tools each agent can use), controlling it (governed access to data and systems, verification of each deliverable), evaluating it, and running it (inference cost optimization, implementation and configuration support). A system of this type goes far beyond a light wrapper around a few LLM calls or a single agent loop. It requires real investment in AI and platform engineering, goes significantly farther than the general-purpose tools the frontier labs have developed themselves, and has to be done well, particularly where the system must be validated for regulated workflows.9 Most buyers have the data, including the completed work an evaluation would be graded against, but not the talent: while a stretched IT team might be able to wire up a workflow, they probably don’t have the expertise in AI engineering and evaluation that deploying to production requires. Outside of specific instances, that expertise is something most AI buyers don’t need and shouldn’t try to build in-house. The work itself is the kind buyers have always paid software vendors to take off their hands.

Why This Is a New Beginning

While these opportunities are distinct, each becomes more valuable when the other exists. An agentic service executes a bounded workflow across several systems and needs somewhere governed to deliver its work: a system of record that can stage a proposal, show a reviewer what would change and why, and record who accepted it. A system of record built for both human and AI work, in turn, provides the most value when there is an agentic service actually capable of doing meaningful work across it. Together they keep accountability where it already sits, with the service responsible for executing and verifying the work, the system of record preserving the source of truth and its audit trail, and the human reviewer making the decisions that carry regulatory weight.

The two do not have to come from the same vendor, provided they share a conceptual ‘contract’: documented endpoints for proposals, traceable review, and deliverables defined well enough to be evaluated. Both are largely unbuilt today. biopharma companies need them and most are not positioned to build them, which is why we see AI as a new beginning for SaaS in this industry, and the opportunity is still open.


Josh is co-founder of Kynetyk, where he writes about AI, builds products at the intersection of AI and human experience, and helps companies design AI strategies that actually scale. Reach out at josh@kynetyk.ai.

  1. The idea predates 2026: in December 2024 Satya Nadella predicted that business applications would collapse into agents (CIO). It peaked in February 2026, when a sell-off in enterprise software stocks was attributed to investor fears that tools like Claude would make SaaS companies obsolete (Fortune). By March, Andreessen Horowitz was describing the “SaaSpocalypse” as market consensus, including the notion that enterprises would vibe code replacements for their internal tooling, while arguing against it (a16z). By June, Meritech was reporting one of the sharpest rebounds in public software in a decade, although a narrow one (Meritech). At the same time, mainstream use of AI is growing, with Pew reporting in June 2026 that about half of US adults report using AI chatbots (Pew Research Center) and OpenAI’s own analysis of ChatGPT Enterprise finding that output grew roughly sevenfold between June 2025 and March 2026 (OpenAI working paper, August 2026), which has also meant that more and more organizations are realizing that AI costs something and that those costs need to, and can be managed. ↩

  2. Where to Buy, Where to Build: An AI Investment Map for Clinical Development maps where AI pays off in clinical trial execution: what sponsors can buy today, what current AI could reach, and the gap between them. ↩

  3. We are not the only ones who see it this way, although others arrive from different directions. An NBER survey of nearly 6,000 executives in the US, UK, Germany and Australia found that 69% of firms use AI and that about nine in ten report no effect on productivity or employment over the prior three years (Yotzov et al., 2026). An INSEAD working paper points to one reason: most studies measure an individual’s output, while organizations do most of their work through patterned collaboration. Its randomized field experiment across 42 teams found productivity gains from an assistant customized with the firm’s own knowledge (Büchsenschuss et al., 2026). On the capacity to build, respondents to CIO.com’s 2026 State of the CIO survey named lack of in-house talent as the top obstacle to implementing AI strategy (CIO). Ethan Mollick makes the related point that AI gains for individuals do not translate into gains for the organization without deliberate change (One Useful Thing). Sequoia and Foundation Capital argue that the opportunity lies in products that take responsibility for the outcome, where a tool leaves that responsibility with the professional using it (Sequoia, Foundation Capital). ↩

  4. METR’s Frontier Risk Report, an external assessment conducted with access to non-public information from Anthropic, Google, Meta and OpenAI, reports that Claude Mythos Preview discovered thousands of vulnerabilities in widely used software, including Firefox and Linux, nearly autonomously (METR, May 2026). The UK AI Security Institute’s independent evaluation concluded that the model can exploit systems with a weak security posture and that more models with these capabilities are likely (AISI). In May 2026 the UK National Cyber Security Centre told organizations to prepare for a “vulnerability patch wave” across open source, commercial, proprietary and SaaS software (NCSC). In July and August 2026, OpenAI, Anthropic and Meta each disclosed that a model had reached a real external organization’s production systems from inside an evaluation environment it believed was isolated; the Cloud Security Alliance’s research note covers the three disclosures together. The labs’ own accounts include Anthropic’s report that Mythos Preview autonomously identified and exploited a 17-year-old remote code execution vulnerability in FreeBSD (Anthropic) and its November 2025 disclosure of an espionage campaign in which AI performed an estimated 80 to 90 percent of the work (Anthropic). The inference that this raises the value of a vendor who owns patching is ours. ↩

  5. The Stanford AI Index reports that 88% of organizations used AI in at least one business function in 2025, up from 78% in 2024 (AI Index 2026, chapter 4). Among US firms that use AI, the Census Bureau finds sales and marketing to be the most common function (52%), ahead of IT (41%) (Census working paper CES-26-25). OpenAI’s analysis of ChatGPT Enterprise (footnote 1) points the same way: engineering and technical staff account for about 11% of weekly active users at the average firm, and growth was “driven primarily by non-Codex output”. Agent use specifically is at an earlier stage: the same AI Index chapter reports scaled agent use in the single digits for nearly all functions. ↩

  6. Connecting an agent to a system of record is technically easy, whether through the Model Context Protocol (MCP), the command line, or custom tools. The MCP specification describes tools as “model-controlled”, meaning the model can discover and invoke them automatically, and it leaves human confirmation and audit logging as recommendations to the client application (MCP specification). Vendor MCP servers generally act with the connecting user’s existing permissions (Atlassian’s, for example). The result is often direct mutation of data in ways that are not easily trackable, auditable or explainable. OWASP classifies the risk as Excessive Agency and recommends human approval of high-impact actions before they are taken (OWASP). In regulated industries the consequences reach beyond the system itself, because a record changed this way can go on to inform a regulatory decision. FDA expects such data to be attributable, legible, contemporaneously recorded, original and accurate (FDA data integrity guidance), and 21 CFR 11.10(e) requires secure, time-stamped audit trails that independently record the entries and actions that create, modify or delete electronic records (eCFR). ↩

  7. Others report the same pattern. METR, whose time-horizon measure is often cited as evidence of long-running autonomy, cautions that it is not the length of time an AI can work independently (METR). Its May 2026 report puts the frontier at roughly 12 hours for tasks completed half the time and roughly 1.5 hours for tasks completed 80% of the time (METR). A Princeton study of 15 models found that recent capability gains have produced only small improvements in reliability, including consistency across repeated runs (Rabanser et al.). On context, Chroma’s test of 18 models found performance becoming less reliable as input length grows (Context Rot). Larger context windows do not remove the problem, because long sessions are repeatedly summarized (compacted) and there is a limit to how much of that a system can absorb: one study found that Claude Code’s compaction preserved 53% of safety rules after one round and 10% after five (Zerhoudi et al.). Persistent memory is the usual remedy, and it is generally written and maintained by the agent itself. On a benchmark of memories invalidated by later events, the best model tested reached 55.2% accuracy (Chao et al.), and a study of small open-weight models found the larger ones more likely to act on a stale note made to look current (Hu and Ramachandran). Curated context files are no sure fix either: across several agents and models, repository context files did not generally improve task success and raised inference cost by more than 20% (Gloaguen et al.). In our experience the informal version of this is worse. Long-running agents with filesystem access produce large numbers of markdown notes, plans and summaries that go stale quickly and are later read back as if current, and a coworker’s value depends on reliably remembering far more than the last 12 hours of work. Anthropic’s engineers write that models tend to lose coherence on lengthy tasks as the context window fills, and describe moving to separate planner, generator and evaluator agents in response (Anthropic). ↩

  8. The human-factors literature anticipated this. Bainbridge’s “Ironies of Automation” describes the monitoring role as “very boring but very responsible” and notes that a job “deskilled” by being reduced to monitoring “is difficult for the individuals involved to come to terms with” (Bainbridge, 1983). Recent studies of coding agents describe the same work. Interviews with 17 experienced developers found oversight spread across control before a run, co-planning, real-time monitoring and review afterwards, with agent-generated code difficult to review (Dhanorkar et al.), and a longitudinal study of 802 developers at one company found that per-reviewer workload roughly doubled as output doubled (He et al.). Closer to our field, a 2026 survey of 637 scientists found that about 46% of those who saved time with AI spent more than a quarter of the time saved auditing and verifying its output (Codreanu et al.). Anthropic’s study of its own engineers reports the same from inside a lab, describing a “paradox of supervision” in which using Claude effectively requires supervision, and supervising it requires the coding skills that heavy AI use may erode. More than half of those surveyed said they could fully delegate only 0 to 20% of their work (Anthropic, December 2025). The picture is not one-sided: a Stanford study found workers positive about agent automation for 46% of tasks, while preferring more human agency than experts consider technically necessary (Shao et al.). ↩

  9. These tools are sophisticated pieces of engineering. A source-code study of eleven production coding harnesses, Claude Code, Codex CLI and Gemini CLI among them, covers roughly four million lines of code and describes the harness as the runtime that couples a model to the world through a loop, tools, context management, safety controls and orchestration (Barbaste et al.). The labs’ own documentation agrees: Anthropic describes Claude Code as an agentic harness around the model that provides tools and manages context, with subagents, checkpoints, permission modes and sandboxing (Claude Code documentation), and its Managed Agents architecture separates a session log, a harness loop and a sandboxed execution environment (Anthropic). OpenAI’s Codex, Google’s Gemini CLI and Microsoft’s Copilot Studio occupy similar ground. Driven by a capable operator, they do a great many things well. What they do not yet offer is the predictability a business would need to let them run unattended on a specific task and count on acceptable work product. Anthropic’s own guidance is “If you can’t verify it, don’t ship it” (best practices), and Microsoft’s is to keep a human in the loop for high-stakes tasks. On the Remote Labor Index, which grades agents on real freelance projects, the best-scoring agent produces a deliverable judged at least as good as the human standard on about one project in five. See also the METR figures in footnote 7. ↩