Home · Blog

Blog

My Articles and Advice

Field notes from production — the patterns that hold up, the mistakes worth avoiding, and the checklists I actually use.

October 31, 2026 · 6 min read

Landing Zones, Explained Simply

A landing zone is the pre-built foundation every workload in your cloud stands on: how accounts are organized, how networks connect, who can do what, and which guardrails nobody can switch off. Get it right once and every project after it inherits sane defaults.

The smallest setup that still counts

  • One management account, plus separate accounts for shared services, logs, and each workload family.
  • Identity federation to your existing directory — no long-lived IAM users, ever.
  • A hub network with inspected egress, so nothing reaches the internet by accident.
  • Preventive guardrails (region locks, encryption requirements) before detective ones.

The mistake I see most

Teams build one giant account "to keep it simple" and pay for it for years in blast radius and billing fog. Splitting accounts early is nearly free; splitting them late is a migration project of its own.

Takeaway: three accounts, federated identity, and inspected egress beat any amount of tooling added later.

October 31, 2026 · 8 min read

5 Terraform Habits That Prevent Outages

Outages from Terraform are rarely about the tool — they are about ceremony. These five habits are the difference between calm teams and 2 a.m. pages.

  1. Pin everything. Providers, modules, and Terraform itself. A surprise minor version is how plans change under you.
  2. Plan on every PR, apply on merge. The plan output is the review — screenshots of it belong in the PR.
  3. Lock remote state. DynamoDB locking or your backend's equivalent. Two applies at once is a corruption story.
  4. Small blast radius per workspace. If one apply can touch the network, the database, and DNS, split it.
  5. Import, don't recreate. ClickOps resources get imported with moved blocks and a test plan — never rebuilt live.

Takeaway: boring Terraform — pinned, planned, locked, and small — is the most reliable infrastructure you will ever run.

November 28, 2025 · 5 min read

Cut Your AWS Bill Without Slowing Down

Almost every audit I run finds the same low-risk wins. None of them require architecture changes, and together they usually clear 25–35%.

  • Rightsize first. Anything averaging under 20% CPU for a month drops a size. No exceptions, no meetings.
  • Graviton where it fits. Most stateless services move with a rebuild and run ~20% cheaper.
  • Lifecycle your storage. S3 Intelligent-Tiering plus 90-day transitions clear the forgotten-bucket tax.
  • Kill the zombies. Unattached EBS, idle load balancers, old snapshots — tag owners, wait two weeks, delete.

Takeaway: want the full list for your accounts? The infrastructure audit finds every one of these in five days.

Discuss your cloud