How We Cut AI Code Hallucinations by 80% in Mobile Code Generation

SA

By Suraj Ahmed

29th Sep 2026

Last updated: 29th Sep 2026

How We Cut AI Code Hallucinations by 80% in Mobile Code Generation

A 2024 empirical study of six popular LLMs generating code across 10,000+ tasks found that 19.4% of all generations contained hallucinations — invented APIs, non-existent packages, syntactically valid nonsense that compiles but does nothing. In more complex agentic settings, the same class of bug climbs past 40%. For a text-completion demo, that's a curiosity. For an AI mobile app builder that ships a compilable React Native app in under a minute, every hallucination is a broken preview, an angry Slack ping, or a customer who never comes back.

We spent most of 2025 rebuilding the code-generation pipeline behind RapidNative around a single rule: the model gets to propose. The system decides whether it commits. That inversion cut AI code hallucinations in our production pipeline by roughly 80% — measured as the difference between the raw model's rejected proposals and what ends up on disk after the guardrails run. This post walks through the five engineering layers that made it work, with actual file paths and real code from the codebase.

If you're building anything agentic — an AI coding tool, a data pipeline that lets an LLM write SQL, a chatbot that schedules meetings — the same pattern applies. Freeform generation is where the hallucinations live. Constrained, validated, tool-mediated generation is where they die.

Developer debugging code on laptop with dark theme editor A single hallucinated import can break a whole preview. Photo by David Rangel on Unsplash

Why generic hallucination advice fails for mobile code

Before the five layers, a quick note on why the standard playbook — retrieval augmented generation, chain-of-thought prompting, higher-parameter models — didn't move our numbers much on its own.

Mobile code fails in ways web code doesn't. A React web app imports react-router-dom and it works. A React Native app that imports react-native-linear-gradient on Expo Go crashes silently because it needs a config plugin; the correct package is expo-linear-gradient. We had 1,681 projects that made exactly this mistake in a three-week window before we caught it. A model trained on the whole internet has seen react-native-linear-gradient in ten thousand tutorials and one hundred correct-for-Expo articles. Statistics wins.

Schema errors don't show up until runtime, in the wrong place. If the model generates a CREATE TABLE migration that enables Row Level Security without a policy, Postgres accepts it. Then the app boots, every query returns zero rows, and the screen renders empty with no error anywhere. The debugger points at the frontend. The bug is in the migration. Retries don't fix this — the model has no signal that anything went wrong.

Retrying is destructive. Streaming a partial response mid-generation, hitting a timeout, and restarting the request has to be done carefully or the user sees the same paragraph twice. Reliability isn't just correctness — it's correctness without visible chaos.

The generic advice — "add tests," "use RAG," "give it better prompts" — presumes a world where the model's mistakes are recoverable. Ours are not. So we built the recovery in.

Layer 1: Tool-use over freeform code

The first and biggest lever was refusing to let the model write freeform code strings at all.

In the old pipeline, the LLM streamed markdown fenced code blocks; a post-processor parsed them out and wrote files. That is a hallucination factory. The model can invent file paths, produce syntactically-broken JSX, generate two files with conflicting exports, or forget the extension. Every one of those is a real bug we filed.

The new pipeline defines every file operation as a Zod-validated tool, in src/lib/coding-agent/tools.ts. The model doesn't write files. It calls a function. And the function checks the arguments before they take effect:

begin_write_file: tool({
  inputSchema: z.object({
    path: z.string().describe("File path, e.g. 'App.tsx'")
  }),
  execute: async ({ path }) => {
    if (!exists && isScreenPath(path) && !newScreensThisTurn.has(path)) {
      if (newScreensThisTurn.size >= MAX_NEW_SCREENS_PER_TURN) {
        return {
          ok: false,
          error: `Screen limit reached: this response already created ${MAX_NEW_SCREENS_PER_TURN}...`
        };
      }
    }
    // ...
  }
})

Notice the split: begin_write_file locks in the path, and only then can write_file_content follow with the body. That serialization stops a whole class of race-condition hallucinations where the model tries to nest edits.

The hard caps aren't polite suggestions in the system prompt — they're enforced in code. MAX_NEW_SCREENS_PER_TURN = 5. Try to create a sixth screen and the tool rejects it, tells the model why, and the model gets to try again in the same turn. In practice the sixth-screen problem was our #2 cause of "the preview is blank" tickets: the model would over-scaffold, run out of context, and produce truncated files. Capping the scope at the tool level, not the prompt level, made it go away.

The same pattern extends to the database. Instead of a JSON DSL for schema mutations (which the model would happily fabricate), the db_migration_new tool takes a plain SQL string:

db_migration_new: tool({
  description: "Create a new SQL migration. Pass `sql` with the full statement(s)...",
  inputSchema: z.object({
    name: z.string(),
    sql: z.string()
  })
})

Same language the migration will be applied in production. Same syntax the model has seen in tens of thousands of GitHub repos. The result is validation authority: what we test is what ships.

Layer 2: The real-Postgres commit gate

If Layer 1 is about proposing safely, Layer 2 is about never committing broken proposals.

Every SQL migration the model writes runs through a validation gate before the file lands on disk. The gate is a real Postgres database compiled to WebAssembly (PGlite), running in the same Node process as the agent. The migrations from the project so far are loaded, the new migration is applied on top, and only if the result is a valid schema does the file get written.

The commit sequence lives in src/lib/coding-agent/db/sql-tools.ts:

const pg = await session();
const schemaBefore = pg.summary();
const dupsBefore = new Set((await pg.duplicateIds()).map((d) => `${d.table} ${d.id}`));
try {
  await pg.exec(body);
} catch (err: any) {
  cached = null;
  return {
    ok: false,
    error: msg,
    hint: "Fix the SQL and call db_migration_new again."
  };
}

Three things are load-bearing in that snippet:

  1. The migration applies on top of the existing chain, not in isolation. A CREATE TABLE workouts migration that would collide with an earlier workouts table fails here, in the same call, before it can land as a broken file that permanently poisons the project.
  2. On failure, the cached engine is discarded (cached = null). This is subtle but critical: a half-applied schema is a lie, and future validation runs off a lie will lie back. The next call rebuilds from disk.
  3. The model gets the error inline. It doesn't retry blindly; it retries with the exact PostgreSQL error message in its context.

Server room with racks of hardware and blue lighting Real Postgres, running in WASM, in the same process as the agent. Photo by Taylor Vick on Unsplash

This layer used to run on pg-mem — a Postgres subset written in TypeScript. It was faster. It was also permissive. It happily accepted auth.uid() = <text column> in RLS policies, which real Postgres rejects outright because there's no uuid = text operator. One project shipped where the very first migration failed on real Postgres at CREATE POLICY. No table was created. The seed never ran. Every screen was empty. Zero errors surfaced anywhere the agent could see. db_migration_new had reported success.

A validator that is more permissive than the runtime cannot do its job. We swapped pg-mem for PGlite and pinned the requirement in tests. If you take one thing from this post: your validator must be strictly equal to or stricter than production. Anything looser will silently ship broken states.

Layer 3: Lints that reach the model in the same turn

Passing the SQL parser is a necessary condition, not a sufficient one. A migration that "works" in the language-legal sense can still ship an unbootable app. The lint layer in src/lib/coding-agent/db/lints.ts catches the semantic gotchas.

The one that mattered most:

if (info.rlsEnabled && info.policies.length === 0) {
  lints.push({
    level: 'error',
    table,
    message:
      `"${table}" has RLS enabled but no policies, so every query returns zero rows and the app ` +
      `will look broken with no error. Add a policy in the same migration.`
  });
}

This is a five-line check that eliminated an entire support category. RLS with no policy is the silent-deny footgun — the model has read a thousand Supabase tutorials that turn RLS on as the "secure default." Turning it on without adding a policy denies everything. The app renders. The queries succeed. The rows are all invisible.

The full lint set:

  • Reserved SQL keywords for table/column names (order, user, group)
  • Case sensitivity traps (Postgres folds unquoted identifiers to lowercase)
  • Missing primary keys (duplicate inserts become possible)
  • Unindexed foreign keys (performance regressions)
  • Missing updated_at triggers

Each lint returns as structured data in the tool's response. The model sees:

"workouts" has RLS enabled but no policies, so every query returns zero rows...

...and adds the missing policy in the same turn without a human ever seeing the intermediate state. This is the difference between a linter that runs in CI (useful, but too late) and a linter that runs at commit time (transformative). The model doesn't get to make the mistake; if it does, it doesn't get to keep it.

Layer 4: Deterministic prompts with byte-identical caching

The prompt engineering layer is the least sexy and probably the most quietly load-bearing. It solves a class of problem that isn't hallucination in the strict sense — it's drift. Same input, different output, no clear reason.

The Supabase agent's system prompt is defined as a module-level constant in src/lib/coding-agent/system-prompt-supabase.ts and composed once at import time:

const SUPABASE_FULL_SYSTEM_PROMPT_FIRST =
  SUPABASE_BASE +
  '\n\n<pre-loaded-skills>\n' +
  getFirstMessageSkills() +
  '\n</pre-loaded-skills>' +
  SUPABASE_SKILLS_SUFFIX;

Two decisions matter here. First, the prompt is byte-identical across every request. That lets prefix caching on the model provider side hit reliably, which cuts cost and — more importantly — cuts variance. Same tokens in, same probability distribution out. Cached prompts are cheaper and more stable.

Second, we split the prompt into a first-turn and follow-up variant. First turn loads every skill the agent might need. Follow-up turns trim the skill set to the core. This matches the empirical shape of a conversation: the first prompt is exploratory and needs breadth; subsequent prompts are surgical and need speed. Loading the full skill pack on every turn wastes tokens and — because context length correlates with degradation — worsens the model's own judgment.

The prompt also carries the hard caps that Layer 1 enforces in code. Belt and braces: the model is told "hard cap 5 screens," and if it violates that anyway, the tool rejects the sixth. In the harness we run periodically, cap violations are 3-8% of turns; the result on disk is always compliant. That is the point.

Layer 5: The harness — proof against a real filesystem

The last layer is not a runtime defense. It's the test loop that keeps the other four honest.

The harness (scripts/agent-harness/thread.ts) runs the exact production agent — same system prompt, same tools, same model resolved from the ai_agents database table — against a real filesystem in .harness-runs/<run>/. No mocks. No stubs. The only difference between the harness and the production API route is the storage adapter: production writes files into a Lifo VFS backed by a files table; the harness writes to disk. Everything else is bit-identical.

You run it with:

npm run agent:run "Create a fitness tracker app"
npm run agent:run "Add a settings screen" -- --run .harness-runs/fitness-1

Each run produces a harness-run.json with the full timeline: every tool call, every argument, every result, every lint the model saw, every file that ended up on disk. If a migration failed, you see the exact error the model was shown. If the model ignored a lint, you see it. If a fix went in, you see the diff.

This matters because the interesting properties of an AI code generation system are executable — does the migration apply, does RLS scope rows correctly, does the generated types.ts match the schema? A test that assembles its own mock agent to check these things tests the mock. The whole point of the harness is that swapping only the storage adapter proves the real agent works.

We've caught three whole classes of regression with this harness that no unit test would have found:

  1. A refactor that changed the order of tool registration silently broke migration validation (the migration tool depended on init order that changed).
  2. A DeepSeek prompt-caching change made the "always include the pre-loaded skills" invariant sometimes false; the harness caught it because runs would suddenly start proposing non-existent tools.
  3. A NativeWind-migration path deleted src/db/schema.ts on Lovable-converted projects because it thought the file was a duplicate of a JSON source that didn't exist. A dry-run script over the real template surfaced it before it shipped.

Team collaborating around a laptop with code on screen The harness lets any engineer replay a bug the model produced last week, verbatim. Photo by Annie Spratt on Unsplash

The dependency-heal step: fixing hallucinations that slipped through

Even five layers of prevention don't get you to 100%. Some hallucinations only manifest at bundle time — an import of a package that isn't in package.json, or a native module that doesn't work in the target runtime. So we added a healer.

After a file is written, src/lib/coding-agent/dependency-heals.ts scans the imports, cross-checks against package.json, and auto-adds the missing declarations. It also rewrites known-bad packages to their known-good equivalents:

  • react-native-linear-gradient → expo-linear-gradient (the Expo Go crasher mentioned above; ~1,681 projects fixed retroactively)
  • react-native-chart-kit was imported by 235 projects that never declared it; those imports now get an auto-added dependency

The healer isn't a substitute for prevention — it's an admission that the model's training data doesn't perfectly match our runtime, and the delta is deterministic enough to patch. When we ship a new template with a different set of libraries, we update the healer's mapping in the same PR.

What "80%" actually means

The number that titles this post deserves an honest definition. We measure hallucinations at two points:

  1. The proposal rate — how often the model produces a tool call that would break something. Zod validation failures, migration errors, lint errors, missing imports.
  2. The commit rate — how often a broken proposal ends up on disk.

Layer 1 (tool validation), Layer 2 (Postgres validation), and Layer 3 (lints) close the gap between the two. Comparing a "raw" pipeline where the model's output is written directly against the current pipeline where the guardrails run:

  • Broken migrations reaching disk: dropped ~92%.
  • Screens that fail to render on first preview: dropped ~78%.
  • Support tickets tagged "blank screen" or "app crashed on open": dropped ~81% quarter-over-quarter after the PGlite swap.

Averaged across the three, we call it 80%. The exact number is less interesting than the shape: the biggest wins came from making bad outputs uncommittable, not from making the model smarter. The model got smarter on its own — every few months a new frontier release lands and our numbers improve. The guardrails are what compound.

What we'd tell a team building something similar

Three things, in decreasing order of impact.

One: define your commit gate before you tune your prompt. A prompt improvement moves the median. A commit gate moves the tail. If you're building an AI-powered React Native app builder, or an AI SQL agent, or an AI infra tool — the tail is what wakes you up. Fix the tail first.

Two: your validator must be the runtime. The pg-mem → PGlite swap is the single most impactful change we made in 2025. Every hour we spent on that was worth ten hours of prompt tuning. If you're validating with a mock, an approximation, or a rewritten-in-your-favorite-language subset, you will ship states that pass validation and break in production. Guaranteed.

Three: give the model the error and let it fix it in-place. Retries at the HTTP layer are the wrong abstraction. Retries at the tool-call layer, where the model sees the exact rejection reason inline in its next turn, are transformative. This is what people mean when they say "agents" — a system in a feedback loop with its own consequences, not just a prompt-and-response.

Try it and see the guardrails in action

Every project generated on RapidNative runs through this pipeline. The free tier gets you 20 credits, no card required. Give it something schema-heavy — a fitness tracker with users, workouts, exercises, and sets, or a marketplace with listings, orders, and reviews — and watch the migration validation kick in. If you break something on purpose (ask for a table with an order column, or an RLS setup with no policy), you'll see the model catch its own mistake in the same turn.

That is what 80% fewer hallucinations looks like from the user side: the model still tries dumb things sometimes. It just doesn't get to keep them.

Start now

Ready to build your app?

Turn your idea into a production-ready React Native app in minutes.

Free tools to get you started

Questions

Frequently asked questions

What is RapidNative?

RapidNative is an AI-powered mobile app builder. Describe the app you want in plain English and RapidNative generates real, production-ready React Native screens you can preview, edit, and publish to the App Store or Google Play.

Can I export the code?

Yes. RapidNative generates clean React Native and Expo code that you can export at any time. No lock-in, no proprietary format. Hand it to your developers or keep building inside RapidNative.

Is RapidNative free to use?

Yes. You can build apps on the free plan with no credit card required. Paid plans unlock unlimited AI generations, code export, and direct publishing to the App Store and Google Play.

Do I need to know how to code?

No. Most users build apps by describing what they want in plain English. Developers can drop into the code whenever they want more control, but coding is optional.

How long does it take to build an app?

Most users have a working first screen in under a minute. A full MVP usually takes a few hours instead of the weeks or months traditional development requires.