Prompt
How do I set up declarative infrastructure for a new cloud platform?
Latest observation
Setting up declarative infrastructure for a new cloud platform is mostly about choosing the right IaC tool, defining a repeatable structure, and building a safe deployment workflow.
1) Pick your declarative infrastructure tool
Common choices:
- Terraform / OpenTofu: best for multi-cloud and broad provider support
- Pulumi: declarative-ish, but uses general-purpose languages
- Cloud-native tools: e.g. AWS CloudFormation, Azure Bicep, Google Cloud Deployment Manager
If you’re on a “new cloud platform,” start by checking whether it has:
- an official Terraform/OpenTofu provider
- a native CLI/API with good auth support
- support for remote state and locking
2) Model your infrastructure in layers
Split infrastructure into clear layers, for example:
- foundation: org, accounts/projects, networking, IAM, logging
- platform: Kubernetes, databases, queues, shared services
- application: app-specific resources, DNS, load balancers, secrets
This avoids one giant config and makes dependencies clearer.
3) Define your code structure
A common pattern:
infra/
modules/
network/
compute/
database/
envs/
dev/
staging/
prod/
- modules/: reusable building blocks
- envs/: environment-specific values and composition
4) Use remote state and locking
For tools like Terraform/OpenTofu:
- store state in a remote backend
- enable state locking
- separate state per environment or per stack
This prevents collisions and makes collaboration safer.
5) Parameterize everything important
Make these configurable:
- region/zone
- instance sizes
- network CIDRs
- naming prefixes
- tags/labels
- scaling limits
Avoid hardcoding values that change across environments.
6) Manage secrets separately
Do not store secrets directly in IaC files or state if you can avoid it. Use:
- cloud secret managers
- external secret stores
- encrypted variables
- short-lived credentials for deployment
7) Build in validation
Before applying changes:
- format code
- run validation/linting
- check dependency graphs
- run policy checks if available
Typical steps:
fmtvalidateplan- policy-as-code checks
- manual review for production
8) Adopt GitOps or CI/CD
A strong workflow is:
- Developer opens a pull request
- CI runs plan/preview
- Reviewers approve
- Merge triggers apply to the target environment
For production, require:
- code review
- change approval
- protected branches
- automated rollback or disaster recovery plan
9) Start with a minimal stack
For a new cloud platform, begin with:
- identity and access
- network
- logging/monitoring
- compute or container platform
- DNS and certificates
- secret management
Then add services incrementally.
10) Standardize naming and tags
Create conventions for:
- resource names
- labels/tags
- environment names
- ownership metadata
Example:
app-env-region-component- tags like
owner,cost_center,env,managed_by
11) Test infrastructure like code
Use:
- dry runs / plans
- integration tests
- ephemeral test environments
- policy tests
- drift detection
12) Plan for drift and recovery
You want to know:
- what happens if someone changes things manually?
- how do you detect drift?
- how do you restore state after failure?
Keep documentation for:
- break-glass access
- state recovery
- disaster recovery steps
A good starter workflow
- Choose Terraform/OpenTofu or the platform’s native declarative tool
- Create a repo with modules and environment folders
- Set up remote state and locking
- Define networking, IAM, and logging first
- Add CI that runs plan/preview on every PR
- Require approval before apply
- Roll out services in layers
If you want, I can also give you:
- a starter folder structure
- a Terraform/OpenTofu example
- a CI pipeline example
- or a step-by-step setup for a specific cloud platform