How I Validate Environment Variables at Startup With Zod in Next.js
A missing env var broke my checkout in production. Here's how I validate every variable with Zod 4 so bad config fails the build, not the user.
Admin
6 min read·
A few months ago I shipped a deploy that looked perfect. CI was green, the build passed, the preview URL loaded. Then the first user tried to check out and got a 500. The cause was embarrassing: I had renamed STRIPE_SECRET_KEY in my local .env file but not in the hosting dashboard. The code happily read undefined, passed it to the Stripe SDK, and the failure only surfaced when a real request hit that code path.
That incident pushed me to adopt a pattern I now use in every Next.js and Node.js project: validate every environment variable once, at startup, with a schema. If something is missing or malformed, the app refuses to boot and tells me exactly what is wrong. In this post I'll walk through how I set it up with Zod 4, the Next.js-specific gotcha around NEXT_PUBLIC_ variables, and a few lessons from running this in production.
The problem with raw process.env
In Node.js, process.env is typed as a dictionary of string | undefined. That tells you almost nothing. Every read is an unchecked assumption, and those assumptions tend to be scattered across the codebase:
// lib/db.ts
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: Number(process.env.DB_POOL_SIZE), // "" -> 0, undefined -> NaN
});
There are three silent failure modes hiding in those two lines:
- Missing values become undefined and only blow up when the code path runs, which might be hours after deploy.
- Wrong types slip through because everything is a string. Number("") is 0, Number(undefined) is NaN, and "false" is a truthy string.
- Malformed values such as a database URL with a typo or a test API key in production look fine until they are used.
The fix is to move all of those assumptions into one file and make them explicit.
A single, validated env module
Here is the server-side module I drop into src/env.ts. It uses Zod 4, which added top-level string format helpers like z.url(), the z.stringbool() helper for boolean-ish strings, and z.prettifyError() for readable error output.
// src/env.ts
import "server-only";
import { z } from "zod";
const serverSchema = z.object({
NODE_ENV: z.enum(["development", "test", "production"]).default("development"),
DATABASE_URL: z.url(),
DB_POOL_SIZE: z.coerce.number().int().min(1).max(50).default(10),
STRIPE_SECRET_KEY: z.string().startsWith("sk_"),
RESEND_API_KEY: z.string().min(1),
ENABLE_SIGNUPS: z.stringbool().default(false),
});
const parsed = serverSchema.safeParse(process.env);
if (!parsed.success) {
console.error("Invalid environment variables:\n" + z.prettifyError(parsed.error));
throw new Error("Invalid environment variables");
}
export const env = parsed.data;
A few details are doing a lot of work here:
- z.coerce.number() turns the string "20" into the number 20, and the chained .int().min(1).max(50) rejects nonsense like "abc" or "0".
- z.stringbool() parses values like "true", "1", "yes" and "false", "0", "no" into real booleans instead of letting "false" be truthy.
- .startsWith("sk_") is a cheap sanity check that catches someone pasting a publishable key where a secret key belongs.
- import "server-only" makes the build fail if a Client Component ever imports this file, so secrets can't accidentally be bundled for the browser.
When validation fails, z.prettifyError() prints every problem at once rather than just the first one, which is exactly what you want when you're fixing a misconfigured deploy:
✖ Invalid input: expected string, received undefined
→ at STRIPE_SECRET_KEY
✖ Too small: expected number to be >=1
→ at DB_POOL_SIZE
And the rest of the codebase gets fully typed values with zero extra effort, because the type of env is inferred from the schema:
// lib/db.ts
import { env } from "@/env";
const pool = new Pool({
connectionString: env.DATABASE_URL, // string, guaranteed to be a URL
max: env.DB_POOL_SIZE, // number, guaranteed 1-50
});
I also stopped reading process.env anywhere else in the app. A one-line lint rule (ESLint's built-in no-restricted-properties or n/no-process-env from eslint-plugin-n) keeps it that way, with an exception for the env files themselves.
The Next.js gotcha: NEXT_PUBLIC_ variables are inlined
Server-side variables are read at runtime, so parsing process.env as a whole works fine. Client-side variables are different. Next.js replaces references to process.env.NEXT_PUBLIC_* with literal values at build time. The browser has no process.env object to parse.
The consequence is that inlining only works when you write out the full property access. These forms are not replaced:
// This does NOT work in client code - the values are not inlined:
const { NEXT_PUBLIC_SITE_URL } = process.env;
const key = "NEXT_PUBLIC_SITE_URL";
const url = process.env[key];
So for client variables I build the object by hand, naming each variable explicitly, and validate that:
// src/env.client.ts
import { z } from "zod";
const clientSchema = z.object({
NEXT_PUBLIC_SITE_URL: z.url(),
NEXT_PUBLIC_POSTHOG_KEY: z.string().min(1).optional(),
});
// Each variable must be referenced by its full name so Next.js can inline it.
export const clientEnv = clientSchema.parse({
NEXT_PUBLIC_SITE_URL: process.env.NEXT_PUBLIC_SITE_URL,
NEXT_PUBLIC_POSTHOG_KEY: process.env.NEXT_PUBLIC_POSTHOG_KEY,
});
It's a little repetitive, but the repetition is the point: it's the only shape the bundler can see. Remember too that inlining means these values are frozen into the JavaScript bundle. Changing NEXT_PUBLIC_SITE_URL in your hosting dashboard does nothing until you rebuild, which is another reason I like validating them so the build fails loudly instead of shipping a bundle with undefined baked in.
Failing at build time, not at request time
A schema only helps if it runs early. Because modules execute when they are first imported, the validation runs whenever something imports env. To make sure it runs on every build and every server start, I import the env module from next.config.ts. A broken configuration then stops next build in CI and in the deploy pipeline, long before any user can hit it.
There's one practical wrinkle: some CI steps, like linting or building a Docker image layer, legitimately run without real secrets. For those I add an escape hatch rather than sprinkling fake values everywhere:
const serverSchema = z.object({
// ...
SKIP_ENV_VALIDATION: z.stringbool().default(false),
});
if (process.env.SKIP_ENV_VALIDATION !== "true") {
// run safeParse and throw as before
}
I keep this flag out of production configuration entirely and only set it in the specific CI jobs that need it. If you find yourself setting it in a real environment, that's a sign the schema is wrong, not that validation should be skipped.
Should you use a library instead?
Packages like T3 Env wrap exactly this pattern, including the server/client split and the skip flag, and they work well. I started with the hand-rolled version because it's about 30 lines, has no dependency beyond Zod, and makes the mechanics obvious. If your team works across many apps or wants a shared convention, a library is a reasonable upgrade. Either way, the important decision is the same: environment variables get validated in one place, with a schema, before the app serves traffic.
Lessons from running this in production
- Validate format, not just presence. Most of my real bugs were values that existed but were wrong: a staging database URL in production, a test key in live mode, a port number with a trailing space.
- Prefer defaults only for genuinely optional settings. A default for DB_POOL_SIZE is fine. A default for DATABASE_URL is a disaster waiting to happen.
- Never log the values. z.prettifyError() reports paths and messages, which is safe. Don't add console.log(process.env) while debugging and forget to remove it.
- Keep a committed .env.example that mirrors the schema, so new contributors know what to set. The schema is the source of truth; the example file is documentation.
Since adopting this pattern, I haven't had a single deploy fail because of a missing or mistyped variable at request time. They still happen occasionally, but now they fail in CI with a clear message, which is exactly where I want them.
Key takeaways
Written by Admin
Published October 6, 2026 · Updated Oct 9, 2026