The Real Role of .env
A .env file is how "config comes from the environment" gets practiced: the code never reads files, only `process.env`, and something external injects it. .env is a local development convenience that loads a file into the environment at startup. It must never be committed — once committed, the secret lives permanently in Git history and deleting the file does not remove it.
The community pattern is committing `.env.example` with keys but no values, which new developers copy to `.env`. That works, with two conditions: `.gitignore` must cover `.env` plus every `.env.*` variant (`.env.local`, `.env.production`), and CI must never generate real config by copying the example — that would hardcode secrets into the pipeline definition.
A safer model merges local sources by priority: system environment variables > project `.env.local` > project `.env` > global `~/.env`. Then a temporary override on one machine never pollutes the repo default.
# .gitignore —— 必须覆盖所有变体
.env
.env.*
!.env.example
# .env.example —— 只有键名,没有值
DATABASE_URL=
REDIS_URL=
SESSION_SECRET=
STRIPE_SECRET_KEY=
FEATURE_NEW_CHECKOUT=false
# 本地启动:dotenv 只负责把文件读进环境
import "dotenv/config"
const sessionSecret = process.env.SESSION_SECRET
if (!sessionSecret) {
throw new Error("SESSION_SECRET is required")
}
console.log(sessionSecret.length >= 32 ? "secret ok" : "secret too short")Four Stages: Choose by Team Size
Stage one, plain .env: 1-5 people, no compliance requirement, few secret types (a database password, one third-party key). The critical work here is making startup config validation a hard failure — exit at boot naming the missing key, instead of failing later on the first database call.
Stage two, platform environment variables: managed platforms (Cloud Run, Heroku, Vercel) provide a config UI, suitable for static values (API endpoints, public client IDs). Its weaknesses are no versioning, no approval flow, no audit trail — and every config edit is an implicit release.
Stage three, container injection: Kubernetes puts non-sensitive values in ConfigMaps and secrets in Secrets, injected into pods as env vars or mounted files. Config now travels with the deployment spec and can be reviewed in a dry-run; the cost is that base64 is not encryption, so you still need etcd encryption at rest or an external secret system.
Stage four, a secret manager (AWS Secrets Manager, GCP Secret Manager, HashiCorp Vault): versioning, automated rotation, fine-grained IAM and full audit trails. It fits scenarios with many secrets, frequent rotation needs and compliance demands. The tradeoffs are added network latency, an availability dependency and new friction for local development — which is why it usually arrives last.
# 第一级:启动期硬校验 + 明确报错
const required = ["DATABASE_URL", "SESSION_SECRET", "STRIPE_SECRET_KEY"] as const
const missing = required.filter((k) => !process.env[k])
if (missing.length > 0) {
console.error("Missing required config:", missing.join(", "))
process.exit(1)
}
# 第三级:K8s 注入,配置与部署描述同源
example@prod:
containers:
- name: app
env:
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: app-secrets
key: database-url
- name: LOG_LEVEL
value: "info"
envFrom:
- configMapRef:
name: app-configSecret Rotation: Overlap Windows Are the Basic Skill
The hard part of rotation is never generating a new value; it is making old and new simultaneously valid during an overlap window. Any rotation scheme must answer: when does the old key die, who triggers it, and how do you roll back. Rotation without an overlap window is a planned brief outage for the whole site.
The standard approach is dual-key support: two environment variables, first ship a version accepting both, then switch to the new key, observe, then revoke the old one. This matters most for symmetric session secrets, where you must decrypt historical tokens with either key, or embed a key id in the ciphertext so each version resolves to its key.
Do not forget hard-to-switch cases such as database passwords. The safe pattern is "create a second user or second password, shift traffic gradually, revoke the old user" rather than changing the password in place. Write the rotation action into a runbook with trigger conditions, owner, verification commands and rollback steps.
# 双密钥重叠:先兼容,再切换,最后撤销
const legacyKey = process.env.SESSION_SECRET_LEGACY
const currentKey = process.env.SESSION_SECRET_CURRENT
function sign(payload: string): string {
return hmac(payload, currentKey!) // 签发只用新密钥
}
function verify(token: string): boolean {
const payload = decode(token)
if (hmacVerify(payload, currentKey!)) return true
return legacyKey ? hmacVerify(payload, legacyKey) : false // 过渡期仍接受旧密钥
}
# 观察期结束后撤销:删除 SESSION_SECRET_LEGACY,代码里的分支自然失效Common Failure Modes and How to Avoid Them
The first failure mode is committing `.env.production`. Block it at two layers with pre-commit and server-side pre-receive hooks, and add a CI step that scans full history with a tool such as gitleaks rather than only the HEAD commit.
The second is mixing config and secrets in one ConfigMap. If logs dump the whole environment, secrets leak. Enforce an allowlist at the code level (log only `LOG_LEVEL`, `APP_VERSION`) and apply the same redaction to CI logs.
The third is local development friction. Secret managers usually require login or VPN; the fix is to cache a short-lived credential locally (`~/.aws/credentials`, or `gcloud auth application-default login`) instead of copying production secrets into `.env`. The fourth is config drift when environments carry different key sets — cross-validate them against one schema (JSON Schema or zod).
// 只允许白名单配置进入日志
const LOGGABLE_KEYS = ["LOG_LEVEL", "APP_VERSION", "NODE_ENV", "REGION"] as const
function safeEnvSnapshot(): Record<string, string> {
return Object.fromEntries(LOGGABLE_KEYS.filter((k) => process.env[k]).map((k) => [k, process.env[k]!]))
}
console.log("startup", safeEnvSnapshot())
# CI 中扫描历史泄露
gitleaks detect --source . --redact --exit-code 1