AI CODE COMMON BUGS risograph thumbnail with rubber-stamped SQLi IDOR XSS SECRET warnings on cream paper, 2026 AUDIT badge

AI-Generated Code Security Bugs: A Vibe Coding Case Study

Saturday I had the assistant vibe-code ScanBoard, a small UI for scan results. Semgrep, gitleaks, nuclei, a Burp pass if I felt like it. Search the findings. Open one by id. Read an artifact. Pull a remote tool report. Login, because I told it the board was multi-user.

It looked done before dinner. Sunday I read the helpers the way I would read someone else's bounty target. The AI-generated code that was supposed to show security bugs had a pile of its own.

Risograph-style desk scene where a laptop spews stamped bugs labeled SQLi, XSS, IDOR, SSRF, and HARDCODED KEY in olive and rust ink
The findings UI looked ready. The backend helpers were not.

What matters

  • The page rendered. Search, the file read, login, and the dependency list were still wrong.
  • ScanBoard shipped classic AI-generated code security bugs: f-string SQL, a finding fetch with no owner check, a hardcoded token, MD5 passwords, a path escape, an HTML echo, an open URL fetch, and a package name that was not on PyPI.
  • What actually helped: bind the SQL, check who owns the id, get secrets out of the file, encode HTML, jail the path, allowlist the URL, then confirm every new dependency exists on the real registry.
Checklist · AI code failures and supply chain

Critical first, then high, then medium, then low. Supply chain is at the top because that is what I miss when I only hunt injection.

SevWhat failsWhat it looks likeWhy the assistant ships it
CHallucinated packages (slopsquatting)import aws-helper-sdk / invented PyPI/npm namesModel invents a package name. Someone registers it before you do.
CUnverified AI-suggested dependenciespip/npm install whatever the chat namedI paste the install line and never open the project page.
Cf-string / concat SQLexecute(f"… LIKE '%{q}%'")Dynamic filters still come out as string glue
CCommand injection via shell helpersos.system / subprocess shell=TrueAI reaches for shell when a library call would do
CSSTI from dynamic templatesrender_template_string(user_html)Flexible page feature becomes RCE-class risk
CInsecure deserializationpickle.loads / yaml.load on uploadsSave the object the short way, and pickle or yaml.load is what comes back.
CPath traversal on file/artifact readsopen(os.path.join(base, user_path))Default upload/download helpers trust ../
CSSRF from URL preview / fetch helpersrequests.get(user_url) with no allowlistI keep missing this until a preview or webhook helper shows up.
CXSS with no output encodingEchoes request data into HTML/markdownIt cannot tell if that string is still user input by the time it hits HTML.
CPlaceholder secrets that shipSECRET_KEY = "sk_live_…" / admin123Demo values look real enough to commit
HOutdated / vulnerable componentsAncient versions from training cutoffThe suggestion is old enough to have a CVE, and it still installs.
HUnpinned or floating versionspackage>=1 / latest tags in AI snippetsLockfile never created; next install can swap the payload
HRisky install / postinstall hooksNew deps with preinstall scriptsThe new package runs a script on install. I did not read it.
HTyposquat-adjacent suggestionsNames one char off real libsLooks legitimate in a diff; resolves to the wrong publisher
HHardcoded JWT / API tokensjwt.encode(…, "my-secret-key")Works in the demo, forge-any-user in prod
Hget_by_id with no ownership checkSELECT … WHERE id=? then return itPrompt asked for CRUD, not authz
HAuth middleware assumed, never wiredRoutes work anonymously in the scaffold"Add auth later" that never lands
HDEBUG / verbose mode left onapp.run(debug=True) / Django DEBUGScaffold defaults that never flip for prod
HMD5 / SHA1 "just hash the password"hashlib.md5(password.encode())Prompt said hash; model picks the shortest one
HHappy-path business logic onlyRefunds, coupons, quotas, racesImplements the demo story, not the abuse cases
MLog injectionlogger.info(f"user={username}")Logging looks harmless, so the raw username or query goes straight in.
MNo CSRF on cookie-auth formsState-changing POST with session cookie onlyCookie login, then a POST with no token. Easy to skip.
MCORS wide openAccess-Control-Allow-Origin: *AI "makes the frontend work" by disabling origin checks
MClient-only validation / mass assignmentChecks in UI; user.update(**request.json)Server trusts role/is_admin fields from the client
MSibling endpoints left insecureOne route fixed, the next still concatsAI patches the file you pointed at, not the copies
LSecurity headers never setNo CSP, no frame denial, default server bannerThe scaffold runs. Headers are not the feature, so they never show up.

Why AI-generated code keeps failing review

The model is trying to finish the prompt. "Search findings" becomes an f-string because that reads like the sentence I typed. "Open finding by id" becomes a SELECT with a placeholder and no idea who is allowed to see it. "Fetch the tool report" becomes requests.get(url). Nobody asked it who owns the row, so it did not invent an owner check.

On these helpers the burn was HTML. The title is still the user's string when it lands in the page. Logs are the same miss on a later prompt, when the raw query gets printed because logging feels harmless. SQL has been a bit cleaner lately. I still read every dynamic filter. Say "hash the password" and you can get MD5. That was the short answer, so that is what came back.

This is not just my weekend luck. Veracode's GenAI Code Security research keeps landing near the same number: only about half of AI code-generation tasks produce secure code when you leave security out of the prompt. Roughly 45% introduce a known flaw class such as SQL injection, XSS, log injection, or weak crypto. Newer models write code that compiles more often. The security pass rate has stayed stubbornly flat.

How ScanBoard maps to OWASP Top 10 2025

If you are hunting by framework, start with the current OWASP Top 10 2025. Supply chain is a first-class category now, which matches how vibe coding actually fails: injection in the helpers, then fake packages in the requirements file.

Map · OWASP Top 10 2025 on ScanBoard

Not a full audit. Just where the weekend helpers landed against the current Top 10, including the new supply-chain category.

OWASP 2025What showed upStatus
A01 Broken Access ControlIDOR on get_finding; open SSRF fetch; path escapeHit
A02 Security MisconfigurationMissing security headers on the scaffoldCommon on vibe apps
A03 Software Supply Chain FailuresTwo hallucinated pip names; unpinned depsHit (reqs file)
A04 Cryptographic FailuresMD5 password hash; hardcoded SECRET_KEYHit
A05 Injectionf-string SQL; XSS title echoHit
A07 Authentication Failuresadmin123 + shared board tokenHit

The weekend build - ScanBoard

The prompt was dull on purpose: search findings, open one by id, log in, read scan artifacts, fetch a remote report. What came back looked like normal Python. I pointed it at SQLite and kept going on the UI. These are the helpers from that pass, unchanged except for file names.

Finding 1 - SQL injection in finding search (CWE-89)

Search was an f-string. It reads like English, which is why it keeps showing up in AI-generated code. Funny place for it: the box that filters security findings.

SB-01 · high · manual · open app/scan_search.py

SQL injection in findings search (CWE-89)

import sqlite3


def search_findings(db_path: str, q: str):
    conn = sqlite3.connect(db_path)
    sql = (
        f"SELECT id, project, title, severity, tool "
        f"FROM findings WHERE title LIKE '%{q}%' OR raw LIKE '%{q}%'"
    )
    rows = conn.execute(sql).fetchall()
    conn.close()
    return rows

Evidence search_findings(db, "%' OR '1'='1") returned every finding row, including other projects.

A normal title returned one row. %' OR '1'='1 returned every finding, other projects included. The fix is a bound parameter. I still had to write that myself.

sql = "SELECT id, project, title, severity, tool FROM findings WHERE title LIKE ?"
rows = conn.execute(sql, (f"%{q}%",)).fetchall()

Finding 2 - IDOR on finding detail (CWE-639)

The id lookup is parameterized. That part is fine. The function does not take a user, and it does not check a session. Whoever can call it gets the row, raw column included. Client-only auth checks are the same miss on BaaS-style scaffolds - the UI hides the button, the API still returns the row.

SB-02 · high · manual · open app/scan_auth.py

Broken access control on /findings/{id} (CWE-639)

def get_finding(db_path: str, finding_id: int):
    conn = sqlite3.connect(db_path)
    row = conn.execute(
        "SELECT id, project, title, severity, owner, raw FROM findings WHERE id=?",
        (finding_id,),
    ).fetchone()
    conn.close()
    return row

Evidence get_finding(3) returned the row with no session and no owner check.

I called get_finding(3) with no session. The row came back. Once that function is a route, another analyst's id comes back the same way. Longer version of the test is in the IDOR hunting playbook.

Finding 3 - hardcoded secret + MD5 passwords (CWE-798, CWE-327)

SB-03 · critical · manual · open app/scan_auth.py

Hardcoded board token and MD5 password hashing (CWE-798 / CWE-327)

import hashlib
import sqlite3

SECRET_KEY = "sk_live_51HxYzScanBoardProdKey"
ADMIN_PASSWORD = "admin123"


def hash_password(password: str) -> str:
    return hashlib.md5(password.encode()).hexdigest()


def login(db_path: str, username: str, password: str):
    ph = hash_password(password)
    conn = sqlite3.connect(db_path)
    row = conn.execute(
        f"SELECT id, username, role FROM users "
        f"WHERE username='{username}' AND password_hash='{ph}'"
    ).fetchone()
    conn.close()
    if not row:
        return None
    return {"token": SECRET_KEY, "user": row}


def get_finding(db_path: str, finding_id: int):
    conn = sqlite3.connect(db_path)
    row = conn.execute(
        "SELECT id, project, title, severity, owner, raw FROM findings WHERE id=?",
        (finding_id,),
    ).fetchone()
    conn.close()
    return row

Evidence Login returned the hardcoded SECRET_KEY as the session token. Password storage used hashlib.md5.

Login hands back SECRET_KEY as the token. That string lives in the file, next to ADMIN_PASSWORD = "admin123". Passwords go through hashlib.md5 because I said "hash the password" and did not name a KDF. The user lookup is another f-string, so the login query is injectable too. If you are hunting keys that actually work, the hygiene in Secrets That Pay is the one I use.

Finding 4 - path traversal, XSS, open fetch (CWE-22, CWE-79, CWE-918)

SB-04 · critical · manual · open app/scan_artifacts.py

Unsafe artifact read, HTML title echo, open URL fetch (CWE-22 / CWE-79 / CWE-918)

import os

import requests


def read_artifact(base_dir: str, path: str) -> str:
    full = os.path.join(base_dir, path)
    with open(full, "r", encoding="utf-8") as handle:
        return handle.read()


def fetch_tool_report(url: str) -> str:
    response = requests.get(url, timeout=3)
    return response.text[:500]


def render_finding_title(title: str) -> str:
    return f"<html><body><h1>Finding: {title}</h1></body></html>"

Evidence read_artifact(artifacts, '../scanner.env') escaped the folder. render_finding_title echoed raw HTML. fetch_tool_report(url) had no host allowlist.

One file, three sinks. os.path.join plus ../scanner.env read a sibling file. The title helper dropped the string into an <h1> with no encoding. fetch_tool_report called requests.get on whatever URL I passed. That last one is the preview feature I keep accepting without an allowlist - I waved it through because the UI looked finished.

Finding 5 - hallucinated dependencies / slopsquatting (CWE-1357)

After the helpers "worked," I asked for an export bridge so ScanBoard could pretty-print Semgrep JSON and nuclei output. The assistant added two new package lines next to requests. Neither name was a real project.

SB-05 · critical · manual · open requirements.txt

Hallucinated PyPI packages in AI-suggested requirements

# AI-suggested extras for ScanBoard report export
requests>=2.28
scanboard-semgrep-bridge==0.3.1
nuclei-result-parser>=1.0

Evidence pip index versions returned no matching distribution for scanboard-semgrep-bridge or nuclei-result-parser. Both names were inventable on PyPI the next morning.

scanboard-semgrep-bridge and nuclei-result-parser looked like every other helper package name. Neither was on PyPI. That is slopsquatting bait: an attacker registers the invented name, waits for the next paste, and owns the install. I never ran pip install on either. The fix is boring and mandatory - open the registry page, confirm the publisher, pin what you keep, and refuse names the model invented. Full hygiene is in the supply chain playbook.

How I use the AI code checklist on a diff

I do not re-read the whole table every time. I start at the top of the diff: new imports, then requirements or package.json, then any string that looks like a key. After that I jump to SQL, HTML, open(, subprocess, and requests. If the assistant added a second copy of a route, I check that copy too. It often "fixes" the one I complained about and leaves the twin alone.

How I verify AI-generated code before merge

Scanners will not clear the PR for me. This is the short pass I still tick by hand when the diff came from an assistant.

ScanBoard · pre-merge gate
AI / vibe-coded PR review

Tick what you verified before merge. Same panel layout I use when tools dump findings into the dashboard and a human still has to clear the gate.

semgrep gitleaks nuclei manual
Sections
5
Checks
21
Blockers
authz · sqli · secrets
Pass rule
all boxes + 1 human pass
manual · burp
Auth and access
4 checks
semgrep · codeql
Injection and output
4 checks
gitleaks · trufflehog
Secrets and crypto
4 checks
nuclei · sca
Files, URLs, deps
5 checks
ci gate
Minimum automation
4 checks

Fast greps I run on AI diffs before the heavier tools:

# fast local greps I run on AI diffs
rg -n "f[\"'].*(SELECT|INSERT|UPDATE|DELETE)|execute\(.*\+|cursor\.execute\(f"
rg -n "innerHTML|dangerouslySetInnerHTML|render_template_string"
rg -n "(api[_-]?key|secret|password|token)\s*=\s*[\"'][^\"']{8,}"
rg -n "hashlib\.(md5|sha1)\(|os\.path\.join\(.*request|requests\.(get|post)\(.*request"
rg -n "eval\(|exec\(|pickle\.loads|yaml\.load\("

How I review vibe-coded PRs now

I stopped asking "is this secure?" The answer is always yes. I also stopped trusting "production-ready" when the assistant says it. That label means the happy path runs, not that authz, encoding, or the registry check passed.

I ask where the request value goes, then I send one bad value at each sink: a quote in the search box, someone else's id, ../ in the file path, a link-local URL in the fetch helper. If those four behave, I read the new dependencies. If they do not, I am not done.

Tool arguments are the same shape of bug when the assistant wires MCP. I covered that in the MCP security testing playbook. Getting the checks into CI, not just into this post, is the laptop to production writeup.

Bottom line

I will keep using the assistant to scaffold. I will not merge AI-generated code until I have read the helpers and the dependency list. The checklist near the top is the pass. ScanBoard is what failed it.

FAQ

Is AI-generated code worse than code I would write?

Line by line, no. It just arrives faster than I review it. The misses I keep seeing are missing ownership checks, HTML and log echoes, placeholder secrets, and packages that do not exist until someone registers them.

What should I check first after a vibe-coding session?

The new imports and requirements, then every object id, then anything that hits SQL, HTML, the filesystem, or an outbound URL. If a secret is a string literal, stop and move it.

Do security prompts fix AI-generated code security issues?

Sometimes the next generation is cleaner. Sometimes the same prompt gives me a bound query, then an f-string ten minutes later. I still read the diff.

What about hallucinated packages and slopsquatting?

Ask for a helper and you can get a package name that is not on PyPI or npm. If you install it anyway, whoever registered that name owns your build. Confirm the project page, pin the version, run SCA, and prefer an SBOM on lockfile changes. More in the supply chain playbook.

Which scanners help on AI PRs?

Semgrep/CodeQL for injection taint, gitleaks or TruffleHog for secrets, SCA for deps. Then still manually probe IDOR - scanners miss a lot of missing ownership checks. See Secrets That Pay and the IDOR playbook.

Can I reuse the ScanBoard snippets?

Copy the checklist and pre-merge gate if you want. Leave the helpers alone - they are the failed versions on purpose.