Config Formats

Terraform State Management in Practice: Remote Backends and State Splitting

The state file is the source of truth in Terraform: it is the collaboration bottleneck and the file most likely to cause a catastrophe. A lost local state means orphaned resources, and concurrent applies mean two engineers acting on the same resource. This guide covers remote backend selection and configuration, the real semantics of state locking including who may unlock and how to unlock safely, correct state splitting by environment and by component, and how import blocks and moved blocks reshape resource addresses without recreating anything.

By LaoHand Team·10 min read·Updated 2026-10-02

Four real risks of a local state file

The first risk is loss. terraform.tfstate is an ordinary file living alongside the working directory, and a failed disk, an accidental delete, or a git clean -fd will destroy it. The real infrastructure then exists only as a list of resources on the cloud provider, and Terraform has lost the mapping from resource to cloud identifier. The next plan either errors that the resource does not exist and attempts to recreate it — usually catastrophic — or reasons from false premises.

The second risk is concurrency. Terraform has no built-in check for who is editing what, and a local state means two people hold independent, mutually unaware copies. Engineer A creates resources and records them in their state; engineer B, without knowing, runs an operation that touches the same class of resources based on the same stale state. The two states diverge permanently. Reconciling them requires operations like terraform state merge, which are easy to get wrong — and a mistake there can delete resources that are still in use.

The third risk is leaked secrets. State files contain database passwords and keys in plaintext. If terraform.tfstate ends up in Git — common when the state directory lives at the repo root and .gitignore was not set up correctly — those secrets live permanently in history, and cleaning them up requires git filter-repo to rewrite the entire repository.

The fourth risk is the absence of an audit trail. Without a central store there is no unified timeline, and questions like "who reduced the production database instance size last Wednesday" cannot be answered except by scrolling chat logs and CI output. Compliance requirements for recorded infrastructure changes end up depending on manual recollection.

The conclusion is direct: a local state file is acceptable for learning, validating provider configuration, and working against local resources such as LocalStack or a kind cluster. Any project touching real cloud resources should configure a remote backend on day one.

# 检查当前是否在用本地 state
ls -la terraform.tfstate terraform.tfstate.backup

# 若 state 曾被提交进 Git,先查历史里是否存在
# git log --oneline --all -- terraform.tfstate

# 从本地 state 迁到远端(S3 backend 为例)
# terraform {
#   backend "s3" {
#     bucket         = "acme-tfstate-prod"
#     key            = "network/terraform.tfstate"
#     region         = "ap-northeast-1"
#     dynamodb_table = "acme-tf-locks"
#     encrypt        = true
#   }
# }
# terraform init -migrate-state -backend-config=backend.hcl

# 确认迁移后本地不再残留
# terraform state list

Choosing and configuring a remote backend

S3 is the most common choice: a mature ecosystem, native support for state versioning and encryption, and locking via a DynamoDB table or S3 conditional writes. The two configuration items that matter most are encrypt = true for server-side encryption and bucket versioning — once versioning is on, a mistakenly overwritten state can be recovered from a prior version, and that is the only thing that rescues an accidental apply.

There are two styles of backend configuration. The backend "s3" block in a .tf file holds only non-sensitive, committable settings, while the concrete bucket, region, and key are passed at init time with -backend-config=xxx.hcl and any secret values are injected through environment variables. This keeps credentials out of Git and stops each engineer from hard-coding a different bucket name.

Terraform Cloud and Terraform Enterprise are the managed option, suited to teams that want state storage, run history, and approval policy handled by a platform. The team then maintains no storage backend at all, at the cost of platform dependency and less flexibility than self-hosted S3 for large-scale local execution. GCS, Azure Blob, and Consul are also officially supported; selection mostly follows the infrastructure and compliance requirements a team already has.

One frequently missed but critical rule: backend blocks cannot reference variables, functions, or any dynamic expression, because Terraform must know where to fetch state before it can evaluate configuration. A common workaround is composing per-environment differences into the key or giving each environment its own prefix within a bucket, injected per environment by -backend-config in CI.

# backend.tf:只写可入库的部分
# terraform {
#   required_version = ">= 1.6.0"
#   backend "s3" {
#     encrypt = true
#   }
# }

# backend.hcl:每个环境一份,进 CI 变量或本地 .gitignore
# bucket         = "acme-tfstate-prod"
# key            = "services/payments/terraform.tfstate"
# region         = "ap-northeast-1"
# dynamodb_table = "acme-tf-locks"
# workspace_key_prefix = "envs"

# 本地开发环境
# terraform init -reconfigure -backend-config=backend.dev.hcl

# CI 中按环境注入(GitHub Actions 示例)
# terraform init -reconfigure \
#   -backend-config="bucket=acme-tfstate-prod" \
#   -backend-config="key=services/${{ env.SERVICE }}/terraform.tfstate" \
#   -backend-config="region=ap-northeast-1" \
#   -backend-config="dynamodb_table=acme-tf-locks"

# 强烈建议:为 state bucket 打开版本控制,这是唯一的误操作恢复途径

State locking: semantics, diagnosis, and safe unlocking

Locking ensures that only one operation at a time — plan, apply, import, state mv — accesses the state file. Note that terraform plan takes the lock by default, which surprises many people: on the S3 backend it is effectively instantaneous, but on a Consul backend or over a poor connection it introduces real latency.

A failure to acquire produces an "Error acquiring the state lock" message. The first move is not to unlock but to determine whether the lock is genuine or stale. A genuine lock means another process is running, and you should go check whether that CI job or terminal is actually alive. A stale lock means the holder crashed or was killed and left its lock entry behind in remote storage.

For the S3 backend with dynamodb_table configured, the lock lives as an item in the lock table, and the Info field carries the holder information; the ID shown in the error message is exactly the value required by force-unlock. Once you have confirmed the lock is stale, terraform force-unlock <LOCK_ID> is the officially supported way to release it.

force-unlock is dangerous because it bypasses the assumption that the holder is dead. If the real holder was merely slow rather than gone, forcibly unlocking lets two operations write state at once and produces a divergence that is very hard to repair. The rule is therefore to verify that the holding CI job or terminal has genuinely terminated before unlocking, and to record the action in the team channel. This is also why many teams add job timeouts and orphan handling to their CI configuration.

# 查看锁状态(命令是即时的,plan/apply 会自动释放)
# terraform force-unlock -force <LOCK_ID>   危险,见下

# S3 + DynamoDB:直接查锁表
# aws dynamodb get-item \
#   --table-name acme-tf-locks \
#   --key '{"LockID": {"S": "acme-tfstate-prod/services/payments/terraform.tfstate-md5"}}' \
#   --projection-expression "Info,ID,Path,Who,Operation,Created"

# 排查顺序:
# 1) 确认持锁的 CI 任务/本地终端是否真的还在跑
# 2) 若进程已终止,读取报错中的 LOCK_ID
# 3) terraform force-unlock <LOCK_ID>
# 4) 在团队频道记录:谁、哪个 LOCK_ID、为什么判断为残留

# plan 不加锁(只读场景,仍会读取状态)
# terraform plan -lock=false -out=tfplan

# apply 时指定预先生成的 plan 文件
# terraform show -no-color tfplan > tfplan.txt
# terraform apply tfplan

Splitting state: by environment and by component are different axes

The first splitting axis is environment: one independent state per environment, either as separate backend keys or as workspaces within a single backend. This prevents change in one environment from polluting plans in another, at the cost of having to redeclare or reference shared resources such as a common VPC or IAM roles in more than one state.

The second axis is component boundary. Splitting infrastructure into separate states for networking, databases, the Kubernetes cluster, and application services is most valuable as the team grows: changing a database parameter no longer plans the entire platform, and review scope shrinks accordingly. It also draws permission boundaries naturally, since write access to one state can be granted only to the team owning that component.

The two axes are orthogonal: workspaces can separate environments inside a single state, or each environment directory can hold several component states. A common choice is to isolate environments by backend key, giving each its own lock and version history, and split components by directory with independent states, so that the two dimensions never contend over the same failure domain. The key thing to remember is that cross-state dependencies must be expressed explicitly, for example a database state reading the VPC state outputs through data "terraform_remote_state".

Splitting has a real cost: relationships between resources go from visible inside Terraform to visible across files, which makes it easier to delete something another state still depends on. Document each component state dependencies in its own README, and back critical output values with prevent_destroy lifecycle rules as a backstop.

# 目录结构:环境用 backend key 隔离,组件用目录拆分
# infra/
#   live/
#     prod/
#       network/backend.hcl   + *.tf
#       database/backend.hcl  + *.tf
#       app/backend.hcl      + *.tf
#     staging/...

# 读取其他 state 的输出(跨 state 依赖必须显式表达)
# data "terraform_remote_state" "network" {
#   backend = "s3"
#   config = {
#     bucket = "acme-tfstate-prod"
#     key    = "network/terraform.tfstate"
#     region = "ap-northeast-1"
#   }
# }
# locals {
#   vpc_id = data.terraform_remote_state.network.outputs.vpc_id
# }

# 关键输出加删除保护
# output "db_subnet_ids" {
#   value       = aws_subnet.main[*].id
#   description = "供应用 state 引用,删除前需确认无下游依赖"
#   lifecycle {
#     prevent_destroy = true
#   }
# }

Reshaping resource addresses: moved and import blocks

When you restructure a directory, Terraform by default concludes that the resource at the old address is gone and the one at the new address is new, so the plan proposes destroy followed by create — unacceptable for a real database. The correct tool is a moved block, which rewrites the address recorded in state and touches no cloud resource at all.

A moved block names from and to resource addresses. Running terraform plan shows a single line reporting that the old address has moved to the new one, with no other changes. After apply, the state entry is re-registered under the new address and the resource itself is completely untouched. This is the standard mechanism for extracting inline resources into a module during a modularization refactor.

Moved blocks support variables and expand across for_each and count, so a whole family of resources can be relocated into a module in one declaration. Their limits are that they can only relocate — never create or destroy — and that they cannot cross state boundaries, because states are isolated from each other.

Import blocks solve a different problem: bringing resources that already exist in the cloud but are absent from Terraform state under management. You add an import block next to the resource block naming the id (the cloud identifier) and to (the destination address). Plan then lists the resources as being imported, and after apply they participate in normal state management and can be changed with ordinary diffs.

Import blocks improve on the older terraform import command in three ways: they are declarative, so they live in version control alongside the code and are reviewable; they are repeatable, since importing an already-imported resource is idempotent; and their outcome is visible during plan. The id field requires a recent provider version — older setups need terraform plan -generate-config-out=generated.tf combined with for_each expansion — and care is needed when composing import with moved blocks and for_each.

# moved 块:只改地址,不动云上资源
# moved {
#   from = aws_instance.web[0]
#   to   = module.web.aws_instance.this[0]
# }

# 批量搬迁:配合 for_each 展开
# moved {
#   from = aws_subnet.public
#   to   = module.network.aws_subnet.public[each.key]
#   for_each = { a = "ap-northeast-1a", b = "ap-northeast-1c" }
# }

# import 块:纳管已存在的云上资源
# resource "aws_s3_bucket" "archive" {
#   bucket = "acme-archive-prod"
# }
# import {
#   id  = "acme-archive-prod"
#   to  = aws_s3_bucket.archive
# }

# 验证迁移是否干净(应无 destroy/create 动作)
# terraform plan -lock=false
# terraform state list | grep -E "^(aws_instance\.web|module\.web)"

Routine operations and disaster recovery

Maintain a team runbook for state operations, covering at least four scenarios with explicit steps: emergency unlocking, including who is authorized and how to verify the holder is dead; backup and rollback, including the concrete commands for restoring from version history; state splitting, including the migration order for adding a component state; and the recovery path for a mistakenly destroyed resource. The value of a runbook is that it removes in-the-moment decision making during an incident.

The backup strategy should be simple and executable. Enabling versioning on the backend is the least effort: every apply produces a new version object, and a mistake can be reverted by restoring a specific version. A periodic timestamped cross-region copy, for example syncing to a second bucket, provides a backstop for regional loss or accidental bucket deletion. Remember that state contains secrets, so encryption and access control on the backup must keep pace.

Splitting state must follow an order: first confirm the target resource list in the original state with terraform state list, then adopt them one by one in the new state using import blocks, and only then move the old addresses with moved blocks or remove them from the old state. Reversing this leaves both states believing they own the resource, and the next apply may delete it.

One more easily missed detail: state records provider versions and schema information. When upgrading a provider major version, run terraform init and terraform plan against a copy of the state first to verify compatibility, then perform the upgrade on the real state, and keep the pre-upgrade version object around so you can roll straight back if anything goes wrong.

# 备份与回滚:利用 backend 版本历史
# aws s3api list-object-versions \
#   --bucket acme-tfstate-prod \
#   --prefix services/payments/terraform.tfstate
# 回滚到指定版本
# aws s3api get-object \
#   --bucket acme-tfstate-prod \
#   --key services/payments/terraform.tfstate \
#   --version-id <VERSION_ID> out/terraform.tfstate
# 确认无误后覆盖回去,并保留一份原件

# 跨区域兜底备份(注意 state 含敏感信息,务必加密)
# aws s3 sync s3://acme-tfstate-prod/ s3://acme-tfstate-backup/ \
#   --exclude "*" --include "*/terraform.tfstate" --sse AES256

# 拆分/纳管前先确认资源归属
# terraform state list
# terraform state show aws_s3_bucket.archive

# 升级 provider 大版本前:先备份,再验证
# terraform state pull > /tmp/state-backup.json
# terraform init -upgrade
# terraform plan -lock=false

Official References

Each command links to its official documentation below, so you can verify the latest usage and read deeper.