Posted in Marketplace, Other, Tech Insights

How to Secure AI-Generated Code in 2026: An Engineering Playbook

Rating:

Subscribe and you will promptly receive new published articles from the blog by mail

AI-generated code compiles almost every time, yet it fails security checks more than 40% of the time. If you want to secure software in 2026, stop relying on prompts to enforce your rules. This AI code security engineering playbook covers what works instead: deterministic CI/CD boundaries, table-driven authorization, strict supply-chain controls, and webhook signature verification.

  • AI code compiles at nearly 100% but fails security tests ~44% of the time. The gap hasn’t narrowed in a year.
  • The real risk is plausible code that passes your tests but quietly breaks system invariants (unscoped queries, missing tenant checks, double payments). The authorization mimicry vulnerability is the classic case: code that looks like it checks permissions and doesn’t.
  • Slopsquatting is real: roughly 5–22% of AI-suggested package names don’t exist, and attackers pre-register those names with malware.
  • The Clinejection incident (Feb 2026) showed that an agent with shell access, fed public input, can leak npm credentials and ship a poisoned release.
  • Security has to live in CI, not in prompts. An instruction like “always verify ownership” is a suggestion. A failing test is a boundary.
  • Every protected action should reduce to: allow = f(subject, action, resource, tenant, state, context), loaded server-side and never trusted from the client.
  • The right unit of trust isn’t “human code” vs. “AI code.” It’s a change that crossed defined boundaries with independent evidence: review, tests, and gates.

Why can’t we just tell the AI to be more careful?

Because instructions aren’t enforcement. You can write “always check tenant ownership” into your system prompt, your CLAUDE.md, your Cursor rules, and it’ll genuinely help. But it’s advice, not a control. GitHub says as much: its own responsible-use documentation for the Copilot coding agent warns that AI code review can miss issues, invent findings that don’t exist, and suggest fixes that are themselves insecure. Human review and real tests are still the job.

An AI-assisted code change never touches just one boundary. It passes through at least seven, one after another:

#
Boundary
What can go wrong there
1
Input
Issues, prompts, logs, repo files, web pages, MCP responses can hide secrets, stale assumptions, or hostile instructions
2
Agent
The assistant reads files, installs packages, runs commands, and edits code – bounded only by whatever permissions it was handed
3
Source
Once committed and reviewed, a change becomes durable – mistakes stick
4
Build
Package managers and build scripts execute third-party code with real CI credentials
5
Artifact
A package or image needs to map back to reviewed source and a known builder
6
Deployment
Production credentials, live data, and migration rights all become reachable
7
Runtime
Users, files, webhooks, APIs, and any LLM tools interact with real business state
ai code change pipeline with security gates

Treat each of those as a checkpoint rather than a formality and most “how did this get through?” postmortems simply stop happening. If you want an AI-generated code security checklist for developers, this table is a decent place to start.

Which AI-generated code failure patterns actually deserve their own test?

Not every bug deserves its own test suite. But a handful of AI-generated failure patterns show up often enough, and hide well enough behind passing tests, that each one earns a test of its own.

Failure mode
What the AI typically writes
Why your happy-path tests miss it
What actually catches it
Context truncation
A correct local function that quietly breaks a global rule
The rule lives somewhere the model never saw
Repo map, scoped tasks, invariant tests, owner review
Plausible-but-fake API
A hallucinated method, header, or package that looks right
Types are often permissive enough to compile anyway
Check the real docs, registry, and version before trusting it
Authorization mimicry
role === “admin” and nothing else
Test fixtures are always pre-authorized
Table-driven deny tests, one central policy
Unsafe fallback
Verification fails → code quietly continues with a default
The exception path never gets exercised
Fail closed; test the timeout/malformed/unavailable cases directly
Retry duplication
A charge, refund, or payout that fires twice on retry
Tests only ever send one request
Idempotency keys, atomic transitions, concurrency tests
Patch overreach
The agent “helpfully” touches workflows, lockfiles, or config too
Reviewer’s eyes are on the file they asked about
Path ownership + mandatory review on sensitive files
Test laundering
The assertion changes to match the bug, not the spec
CI goes green either way
Review what the test actually checks before reviewing the fix
Secret leakage
Real credentials end up in a prompt, fixture, or log
They may never touch tracked source at all
Scoped context, secret stores, redaction, rotation
Review monoculture
Same model writes it, same model reviews it
Shared blind spots agree with each other
Independent tools, a human owner, adversarial tests

The “plausible-but-fake API” row deserves a closer look, because it’s turned into an actual attack vector. Researchers tested 16 code-generating models across more than half a million samples. Commercial models invented non-existent package names in roughly 5.2% of recommendations, and open-source models did it in 21.7% of theirs.

Attackers watch for exactly this. They register the fake package name a model keeps suggesting, load it with malware, and wait. Security researchers call it slopsquatting, and it’s as sneaky as it sounds: the fake name often sounds more plausible than the real one. That’s why dependency review for AI-generated packages shouldn’t be optional.

Before you npm install anything a model suggested, check that it exists in the official registry. Then read the actual lockfile diff. Switching on GitHub dependency review makes that diff hard to skip, because it surfaces every dependency change right in the pull request.

What does a secure generation protocol actually look like day-to-day?

secure ai code generation protocol infographic

Break it into three phases: before the agent runs, while it runs, and before anything ships. Suddenly it stops feeling abstract. Think of the lists below as a secure AI code review checklist for developers, and as the day-to-day core of this AI code security engineering playbook.

Before the agent runs

  • Write the security invariants and abuse cases before implementation starts, not after.
  • Flag sensitive paths up front: auth, payments, workflows, migrations, infra, lockfiles.
  • Give the agent only the repo scope, network access, and credentials the task needs. Nothing extra “just in case.”
  • Use sanitized fixtures. Production secrets, real customer data, and signing keys have no business in a prompt.

While it runs

  • Review commands and dependency installs before approving them, not after they’ve already run. That’s where dependency review for AI-generated packages earns its keep.
  • Treat all external content as untrusted (issue text, web pages, repo docs, MCP output), because any of it can carry hidden instructions.
  • Keep the work on an isolated branch or ephemeral environment.
  • Log the task, the identity that triggered it, tool activity, the diff, and the review decision. That trail won’t make a bad patch safe, but it makes the investigation afterward possible.

Before merge and deploy

  • Read the actual diff, not the agent’s tidy summary of what it did.
  • Review the tests first, and check that nobody quietly weakened an assertion to make things pass.
  • Run formatting, type checks, unit and integration tests, secret scanning, SAST, and dependency review on every change.
  • Require a human code-owner on anything touching identity, payments, workflows, infra, or security policy.
  • Build from the reviewed commit using a locked-down CI identity, and keep build credentials separate from production ones.
  • Ship an immutable artifact through a staged, observable, and reversible process.

A quick aside on what “correct” even means here: none of this replaces good business analysis up front. Half the “AI got it wrong” stories we hear trace back to a spec that never said what “correct” meant for that edge case. The model didn’t fail the requirement. The requirement didn’t exist.

How do you test for the authorization mimicry vulnerability?

Yes, and this is where most teams stop too early. A comment that says “always verify ownership” is a nice thought. A function that returns false when ownership doesn’t check out is a control. Only one of those will ever stop a bad request.

Every protected action should reduce to something like:

allow = f(subject, action, resource, tenant, state, context)

Just as important: never trust the client to tell you the tenant ID, the price, or the role. Load them server-side, from the authenticated context or the database, not from whatever the request claims.

type Actor = { id: string; tenantId: string; roles: string[] };
type Order = {
  id: string;
  tenantId: string;
  buyerId: string;
  sellerId: string;
  state: "pending" | "paid" | "cancelled";
};

export function canCancelOrder(actor: Actor, order: Order): boolean {
  if (actor.tenantId !== order.tenantId || order.state !== "pending") return false;

  const participant = actor.id === order.buyerId || actor.id === order.sellerId;
  return participant || actor.roles.includes("tenant-admin");
}

One catch: a correct policy function doesn’t save you if the route handler pulled the order from an unscoped query somewhere else. That’s the authorization mimicry vulnerability again, only harder to spot, because the check exists and looks right. The policy function and the data-loading code both have to get it right; one clean function can’t stand in for the other.

Then write the test that actually tries to break it:

it.each([
  ["other tenant", otherTenantActor],
  ["unrelated user", unrelatedActor],
])("denies %s", async (_case, actor) => {
  const response = await request(app)
    .post(`/orders/${pendingOrder.id}/cancel`)
    .set(authHeader(actor));

  expect(response.status).toBe(404);
  expect(await orderState(pendingOrder.id)).toBe("pending");
});

On validation and injection: these solve two different problems, so don’t let one stand in for the other. Validate shape and limits with something like Zod, then use parameterized queries for the actual database call:

import { z } from "zod";

const searchSchema = z.object({
  query: z.string().trim().min(1).max(100),
  limit: z.coerce.number().int().min(1).max(50).default(20),
});

const input = searchSchema.parse(req.query);
const result = await db.query(
  `SELECT id, title FROM listings
  WHERE tenant_id = $1 AND title ILIKE $2
  ORDER BY created_at DESC LIMIT $3`,
  [req.auth.tenantId, `%${input.query}%`, input.limit],
);

For anything dynamic (sort fields, column names), use a fixed allowlist and never string interpolation. For XSS, lean on your framework’s auto-escaping, and be deliberate (and rare) about rendering raw HTML.

SSRF is worth designing out entirely if you can. For SSRF attack prevention, a fixed destination allowlist beats “let the server fetch any URL the user gives it” every time. If arbitrary URLs are genuinely required, enforce scheme, hostname, resolved IP range, port, redirect behavior, response size, and timeout at the network egress layer. A single upfront string check won’t survive DNS tricks or redirects.

How do you handle webhooks without breaking webhook signature verification?

Authenticate, then deduplicate. In that order, every time.

Stripe’s own docs are explicit that webhook signature verification needs the unmodified raw request body, the Stripe-Signature header, and your endpoint secret. If a global JSON parser touches the body before your webhook route sees it, verification breaks silently.

app.post(
  "/webhooks/stripe",
  express.raw({ type: "application/json" }),
  async (req, res) => {
    const signature = req.header("stripe-signature");
    if (!signature) return res.sendStatus(400);

    let event: Stripe.Event;
    try {
      event = stripe.webhooks.constructEvent(
        req.body,
        signature,
        process.env.STRIPE_WEBHOOK_SECRET!,
      );
    } catch {
      return res.sendStatus(400);
    }

    await acceptStripeEvent(event);
    return res.sendStatus(200);
  },
);

Webhooks retry by design, so assume any delivery might be a duplicate. Record the provider’s event ID under a unique constraint, in the same transaction as your state change:

CREATE TABLE processed_webhook_events (
  provider text NOT NULL,
  event_id text NOT NULL,
  processed_at timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (provider, event_id)
);

On the outbound side, give every create or update request an idempotency key. Just remember that idempotency stops duplicates; it doesn’t replace authorization or server-side amount checks. You need both. If you’re setting this up from scratch, our payment gateway integration work runs into this class of bug often enough that webhook signature verification is basically the first thing we check.

How do you secure AI-generated code on Sharetribe?

A quick note for marketplace builders, since Sharetribe is a common stack. Its transaction definitions restrict which actor can trigger which state, and a privileged transition goes a step further by requiring server context obtained through a client secret that only your backend holds. For a deeper walkthrough, see our guide on securing AI-generated code in Sharetribe marketplaces.

sharetribe privileged transition flow diagram

That’s real protection, but it doesn’t replace business logic on your side. Your custom backend still has to recalculate prices, line items, commissions, and refund eligibility itself instead of trusting whatever the client sent. Test the nasty cases: a normal token attempting a privileged transition (it should get a clean 403), tampered parameters, an invalid prior state, and repeated or concurrent requests. 

The Sharetribe Web Template ships with a reasonable starting test suite, but it’s a starting point, not full coverage. Check its current testing and CI setup against the version you’re actually running.

What happens when things go wrong: does your code fail safely?

This one’s easy to skip because writing tests for the unhappy path is boring. It matters enough, though, that OWASP’s 2025 Top 10 added a brand-new category for it: A10, Mishandling of Exceptional Conditions. It’s one of only two entirely new categories in that edition, which is a strong signal that “how does this fail?” is finally being treated as a security question and not just a reliability one.

AI-generated code has a specific blind spot here: it tends to optimize hard for the success path and improvise everything else. So test the improvised parts on purpose. Put these on your AI-generated code security checklist for developers: verifier timeouts, transaction rollbacks, partial provider outages, malformed uploads, redelivered queue messages, missing config, and exhausted rate limits.

The rule that matters most: security decisions should fail closed. If your identity verifier goes down, that can’t quietly mean “let them in.” And a payment shouldn’t go through just because a check timed out. That said, “catch every error and deny forever” is just a different kind of outage. Good design lives somewhere between the two extremes.

Can CI actually enforce AI code security?

It has to live in CI. AI assistants generate more code and more dependencies than any reviewer can realistically read line by line, so anything that matters needs to be a machine-enforced gate rather than a note in a review comment or a rule everyone’s supposed to remember.

Pin third-party GitHub Actions to reviewed commit SHAs (tags can move), minimize GITHUB_TOKEN permissions, and keep untrusted PR jobs away from anything that holds real secrets:

name: security-gates

on:
  pull_request:

permissions:
  contents: read

jobs:
  verify:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@<REVIEWED_COMMIT_SHA>
        with:
          persist-credentials: false
      - uses: actions/setup-node@<REVIEWED_COMMIT_SHA>
        with:
          node-version-file: .nvmrc
          cache: npm
      - run: npm ci --ignore-scripts
      - run: npm run format:check
      - run: npm run typecheck
      - run: npm test -- --runInBand
      - run: npm audit --audit-level=high
      - run: npm run build

Swap those SHA placeholders for real, reviewed commits before this goes anywhere near production. And –ignore-scripts helps, but it isn’t a silver bullet. If a dependency genuinely needs install scripts, review them deliberately instead of just turning the flag off.

One caveat: a clean npm audit only tells you there’s no known advisory in the registries it checked. It says nothing about provenance or undiscovered malicious behavior. Treat it as one signal among several, and be careful with –force fixes, which can install outside your declared version ranges without much warning.

Add platform-native secret scanning and a real SAST tool such as CodeQL. Then add GitHub dependency review, which surfaces both direct and transitive changes in a PR and can block on policy violations when it’s set as a required check.

Here’s roughly how those gates should map to blocking rules:

Gate
Pull request
Release
Blocks on
Format, types, unit tests
Yes
Yes
Any failure
Authorization & payment tests
Changed paths
Yes
Any failure
Secret scanning
Yes
Yes
Confirmed active secret or policy match
SAST
Diff/PR
Full
New high-confidence finding at your set severity
Dependency review
Yes
Yes
New vulnerable or prohibited dependency
DAST/API scan
Preview/staging
Yes
Confirmed exploitable release blocker
Artifact provenance
Build
Deploy
Artifact not built by the approved workflow + commit

And if you’re going to suppress a finding instead of fixing it, make it accountable: a finding ID, a reason, a compensating control, an owner, and an expiry date. Baseline your existing backlog separately, too. Otherwise a pile of old findings makes it easy to miss the one new one that matters.

If you’re running a SaaS platform rather than a marketplace, the vendor changes but the questions don’t. What does your CI identity have access to? Does a normal user token ever get treated as trusted server context by mistake?

Who gets to run commands – and does dependency review for AI-generated packages catch it?

A real incident makes all this abstract stuff concrete. In February 2026, security researcher Adnan Khan published the Clinejection disclosure, and it’s worth reading in full if you run any kind of AI agent in CI.

Cline’s issue-triage workflow let any GitHub user trigger the agent (the config literally allowed non-write users via a wildcard) while giving that agent Bash, Write, and Edit tools. A prompt injection hidden in an issue title was enough to get code execution inside the workflow.

From there, shared GitHub Actions caches opened a path straight into nightly release jobs and publishing credentials. Cline confirmed the full chain in its own post-mortem: credential rotation turned out to be incomplete, and someone used the still-live npm token to publish an unauthorized cline@2.3.0 release on February 17.

That particular package just installed the legitimate OpenClaw project and did nothing worse, but the access it exploited could easily have shipped arbitrary code to everyone who installed it.

The lesson isn’t “filter your prompts better.” It’s structural. Never give an agent shell access while it’s processing input from the public internet, and never let a low-trust workflow share a cache or credentials with anything that can publish a release.

For comparison, GitHub’s own cloud coding agent runs in an ephemeral, firewalled sandbox, can’t push straight to your default branch, and doesn’t get organization secrets unless you deliberately hand them over. Those are good defaults, but they protect the platform, not your business logic. That part is still on you.

Whatever platform you’re on, a few non-negotiables belong on your secure AI code review checklist for developers:

  • Agents write to a branch or worktree, never directly to a protected default branch.
  • A human approves before any workflow exposes secrets to an agent-authored PR.
  • Only task-scoped credentials live in the agent’s environment. Production credentials never do.
  • Network egress is denied by default and allowlisted only where the task genuinely needs it.
  • Every MCP server and tool has an owner, an inventory entry, and session-scoped access.

What if the AI runs in your product – and how does SSRF attack prevention apply?

This is a different threat model entirely. If your product has an LLM, RAG pipeline, MCP integration, or user-facing agent, the architecture has to assume the model itself can be manipulated. The OWASP LLM Top 10 covers prompt injection (direct and indirect), sensitive data disclosure, unsafe output handling, excessive agency, supply-chain risk, and unbounded resource consumption.

8 rules worth building around from day one:

  1. Treat the model’s output, and anything it retrieved, as untrusted data.
  2. Enforce identity, tenant, resource, action, and state outside the model, not inside a prompt.
  3. Give it narrow, typed tools: no generic SQL access, no shell, and no unrestricted HTTP. SSRF attack prevention applies here too, so any fetch tool should only reach an allowlist of destinations.
  4. Re-authorize every tool call against the real authenticated user, every time.
  5. Require independent human approval for money movement, deletions, permission changes, or anything published externally.
  6. Validate tool arguments and outputs at the destination (the sink), not just on the way in.
  7. Put hard bounds on requests, tokens, recursion depth, tool calls, time, concurrency, and spend.
  8. Log every decision with privacy-aware redaction, and actually test indirect prompt injection instead of assuming your filters catch it.
secure llm tool call flow diagram

Here’s a quick gut-check for rule #2, to see whether you’ve built it or just told the model to follow it. “Never refund an unauthorized order” typed into a system prompt is a suggestion the model can be talked out of. A refund tool that loads the order through the authenticated tenant, applies deterministic policy code, and opens an approval request instead of trusting whatever the model claims is the real boundary.

So when is an AI-assisted change actually “done”?

Probably not the moment the tests go green. Here’s a definition of done that holds up under a postmortem, and you can use it as your final AI-generated code security checklist for developers:

Requirement
What proves it
Task and security invariants are scoped
Issue or PR acceptance criteria
Agent permissions are minimal
Session, environment, and tool config
Diff is independently reviewed
Human + code-owner approval
Tests cover allowed and denied behavior
Unit and integration results
Exceptional paths are tested
Timeout, malformed input, retry, duplicate, rollback cases
Dependencies and workflows are reviewed
Manifest, lockfile, and Action diff
Secrets are absent, access is scoped
Secret scan + environment policy
Security findings are resolved or owned
CI checks + finding register
Artifact maps to reviewed source
Digest and provenance/signature policy
Deployment is observable and reversible
Rollout, alert, rollback, restore evidence

What’s the takeaway for securing AI-generated code in 2026?

“Code a human wrote” versus “code the AI wrote” was never really the right question. What makes a change trustworthy is whether it crossed your system’s boundaries with independent evidence behind it: real review, real tests, real gates. Who or what typed it first matters much less.

Teams that bake their security invariants into tests and CI policy, instead of prompts and tribal memory, can use AI aggressively and still keep production truth in human hands. That’s the whole point of an AI code security engineering playbook: once the boundaries are enforced and not merely requested, the two aren’t in tension.

Roobykon works with marketplace and SaaS teams on exactly this: architecture and trust-boundary reviews, transaction and payment logic audits, CI and supply-chain hardening, webhook signature verification reviews, and AI-enabled workflow reviews. The deliverable is always a prioritized, executable remediation plan, never a vague “your codebase is AI-secure now” checkbox.

Want a second pair of eyes on your setup?

Get in touch and tell us where AI is already writing code in your product. We’ll tell you, honestly, where the actual risk is, and where it isn’t.

Let's talk
Oleksandr Lyubenko
Oleksandr Lyubenko Founder & CEO

Oleksandr Liubenko is the Founder and CEO of Roobykon Software, a company he built from the ground up in 2011 into one of the leading online marketplace development teams in the industry. With over 25 years of experience across software development and business management – including more than 15 years in executive leadership – Oleksandr brings a rare combination of technical depth and strategic vision to every project.

Recommended articles