The Production Checklist for AI-Built Apps

BlogOctober 8, 2026

You built the app in a weekend. Lovable, Bolt, Replit, Cursor or v0 got you to a demo that works. Then you tried to put real users on it and it stalled around 80%.

That is not bad luck. These tools are tuned to make the demo work. They are much weaker at the parts nobody sees until something goes wrong: who can read which row, what happens when a payment webhook arrives twice, how you find out the app is down.

This is the checklist we run before an app takes real traffic. It is written for founders, but every line is something an engineer can check in an afternoon. Work through it top to bottom. Anything you cannot answer is a launch blocker, not a nice-to-have.

Data and access

This is where AI-built apps fail most often, and where the damage is worst. Examples use Supabase because most of the apps we see are built on it, but the questions apply to any backend.

  • Row Level Security is on for every table in an exposed schema. A table in the public schema without RLS can be read and written by anyone who has your anon key, and that key ships in your front end by design. RLS enabled with no policies denies everything, which is the safe default.
  • Check RLS on tables your AI tool created with SQL. Tables created through SQL or migrations, which is what AI tools usually emit, start with RLS disabled. The dashboard Table Editor enables it by default, so a table that looks fine there can still be wide open if it was created by a migration. Supabase's docs say to enable it yourself for tables created with SQL.
  • Every policy is tested as a real user. The dashboard SQL editor runs as a privileged role that skips RLS, so a policy that works there proves nothing. Sign in as two ordinary accounts and confirm A cannot see B's rows.
  • No service-role key in client code. It bypasses RLS completely. If it is in your front-end bundle, every policy you wrote is decoration. Search the repo and the built bundle. Remember that anything prefixed NEXT_PUBLIC_ or VITE_ is public.
  • Secrets live in environment variables, not the repo. Check the git history too. A key that was committed and later deleted is still exposed. Rotate it.
  • Policies do not trust data the user can edit. In Supabase, user_metadata can be changed by the signed-in user. Put roles in app_metadata or a separate roles table.
  • Views and functions respect RLS. A view runs with its owner's permissions by default and can bypass your policies. On Postgres 15 and later, create it with security_invoker = true.
  • Storage buckets have policies. A public bucket serves every file to anyone with the URL. Invoices, IDs and private uploads need a private bucket and policies on storage.objects.

Auth and permissions

  • Every request is authorized on the server. Hiding a button is not access control. Call the API directly as a user who should be refused and confirm that it refuses.
  • Users can only touch their own records. Change the ID in a URL or request body to another user's. You should get a 403 or 404, not their data.
  • Admin screens and admin actions check the role on the server. A route that is merely not linked is still reachable.
  • Reset links, magic links and invites expire and work once. Old links that still work are a standing back door.
  • Login, signup, password reset and one-time codes are rate limited. Otherwise anyone can guess at them all day.
  • OAuth redirect URLs list only your real domains. Remove localhost and old preview URLs before launch.
  • Signing out and disabling a user actually cut off access. Test it with a second browser.

Payments and webhooks

  • Orders are fulfilled from the webhook, not the success page. The browser redirect can be closed, skipped or faked. The provider's webhook is the source of truth.
  • Webhook signatures are verified. Verify against the raw request body and your endpoint secret. Without that, anyone can post a fake "payment succeeded" to your server.
  • The same event can arrive twice without doing damage. Providers retry and sometimes deliver duplicates. Store the event ID and ignore repeats.
  • Requests that create charges use idempotency keys. A double click or a network retry should never bill someone twice.
  • The unhappy paths work. Declined card, expired card, refund, cancellation, downgrade, failed renewal. The plan stored in your database should always match the provider's.
  • Test keys and live keys are separated. Live keys exist only in the production environment.

Error handling and observability

  • Error tracking is installed on the front end and the back end. Sentry or similar. Without it, you hear about bugs from angry users.
  • Failures show a useful message, not a blank screen. Every async call has a loading, error and empty state, and the app has an error boundary.
  • Errors do not leak internals. Users should never see a stack trace, a SQL error or a key.
  • Important actions are logged and searchable. Payments, deletes and role changes: who, what, when. Passwords, tokens and card data stay out of the logs.
  • An uptime monitor alerts a channel someone reads. You should know the app is down before a customer tells you.
  • Calls to other services have timeouts and retries. Email, payments and AI APIs all fail sometimes. The app should fail gracefully when they do.

Tests

AI tools and demos check the happy path: one user, clean data, everything goes right. When the checking model on our QueueMate build wrote its test cases, only 49 of 554 were happy path. That is about 9%. The other 91% were edge, negative, boundary, security and concurrency cases. Real users live there. We explain the method in how we test AI-generated code before it ships.

Run these against your app before anyone else does.

  • Time zones. Records created near midnight, a user in a different zone from your server, daylight saving changes. Store UTC and convert for display.
  • Empty and invalid input. A blank form, 10,000 characters, emoji, quotes, negative numbers, a pasted script tag.
  • Double submit. Double click the pay and save buttons. Refresh mid-request. Look for duplicate rows.
  • Concurrency. Two people grab the last item at the same moment. Two tabs edit the same record.
  • Roles. Try every action as each kind of user and as a logged-out visitor.
  • Dead ends. A session that expires mid-form, the back button, a deep link opened while logged out, a slow connection.
  • One automated end-to-end test of the money path. Sign up, do the core action, pay. Run it on every deploy.

Performance and cost

  • No N+1 queries. A list that fires one query per row is fine at 10 rows and falls over at 10,000.
  • Indexes on the columns you filter and join on, and pagination on every list. Check the slowest page with real-sized data, not demo data.
  • LLM spend is capped. Set max tokens, per-user limits and a spend limit with the provider, and log tokens per request. A loop or one abusive user can burn a month of budget overnight.
  • Expensive endpoints are rate limited. Anything that calls a model, sends email, uploads files or exports data is a bill waiting to happen.
  • Long jobs run in the background. Do not make a request wait on a task that can take a minute.
  • Pages are fast on a mid-range phone. Compress images and lazy-load what is below the fold.

Deploy and environments

  • Staging is separate from production and has its own database. Do not test on real customer data.
  • Environment variables are documented and set per environment. Keep an .env.example. A missing variable is the classic "works on my machine" failure.
  • Database changes are migrations in version control. Changes made by clicking around a dashboard cannot be reproduced or rolled back.
  • Backups are on, and you have restored one. A backup you have never restored is a hope, not a plan.
  • You can roll back in minutes. Know how to put the previous version back, and keep migrations backward compatible where you can.
  • A merge deploys, and checks run first. Type check, lint and tests should pass before anything reaches production.
  • Staging does not email real customers or get indexed by search engines.

Launch day

  • Domain, HTTPS and redirects work. Pick www or the bare domain and redirect the other.
  • Email lands in the inbox. Set SPF, DKIM and DMARC on your sending domain, then send test messages to Gmail and Outlook and check the spam folder.
  • Analytics and conversion events fire in production. Verify on the live site, not in preview.
  • Privacy policy and terms are linked. If you take payments or serve visitors in the EU or UK, check what else applies to you.
  • The small things are done. Custom 404 page, favicon, page titles, share images, robots and sitemap.
  • You have smoke-tested production with a real signup and a real payment. Refund it afterward.
  • Someone owns the first week. Decide who receives the alerts and who fixes what they find.

How to read your result

Data and access, auth, and payments are the sections where a gap can hurt real people. Treat any unanswered line there as a launch blocker. Gaps in tests, observability and deploys will not leak data, but they decide whether you can keep fixing things after launch without breaking something else.

If you cannot tick most of this list and you would rather not spend the next month on it, this is the work we do. We run a paid two-week trial that audits your app, fixes what blocks launch and hands it over ready for production. See how Vibe Code Rescue works, or see how the two-week finish trial works.

If you want to see how we work on our own products, read how one engineer built a full MVP with a team of AI agents.

Frequently asked questions

Is a vibe-coded app safe to launch?

Not by default. It can be, after a review. The usual problems are not in how the app looks. They are in who can read which data, whether secrets are exposed, and what happens when things go wrong. Work through the data and access, auth and payments sections of this checklist before real users arrive.

How do I know if my Supabase tables are exposed?

Check each table for Row Level Security in the dashboard. Tables without it are flagged in the Table Editor and in the Security Advisor. Then test from outside: call your project's REST API with only the anon key and see what comes back. If you can read data you would not want a stranger to read, so can everyone else.

Where should I start if the whole list looks overwhelming?

Start with data and access, then auth, then payments. Those are the sections where a gap can hurt real people. Leave performance and polish for after the app is safe.

Do I need to rewrite my AI-built app to make it production-ready?

Usually not. Most apps can keep what works and fix what blocks launch. A rewrite makes sense only when the foundation cannot be secured or tested, and a short audit will tell you which case you are in.

Does this checklist only apply to Supabase apps?

No. The Supabase lines are examples. The same questions apply to Firebase, Postgres behind your own API or any other backend: who can read what, where the secrets are, and what happens when a request arrives twice.

Building something with AI?

First Mate is an AI engineer agency. Our product-minded AI engineers scope the problem before they build it. You can start with a rough cost and timeline, or talk to us directly.