Pulumi IaC: How AI Skills Cut Failed Redeploys
A controlled intern study on whether Antigravity Skills reduce failed Pulumi redeploy cycles.
Mike Henken
AI Platform Architect
Escaping the
IaC Abyss
Generic LLMs stumble on IaC: deprecated provider args, state drift, and secret handling. This write-up covers a small intern study on whether a scoped Pulumi SKILL.md improves outcomes.
I started with a simple question: can a scoped skill file beat generic chat for Pulumi Python on AWS?
In IaC, the live cloud state is the source of truth. Code is the intent. Skills should enforce that distinction before they suggest deletes.
A dashboard bot once analyzed a state file, proposed a fix, opened a PR, and ran pulumi up successfully on the first try. That was specialized context, not a general model flex. The intern study asked whether we could bottle that context into a reusable SKILL.md.
We compared Cursor rules-only groups against an Antigravity group carrying a Pulumi architect skill. The skill group spent less wall-clock time waiting on failed redeploys in the exercise window.
Why Generic LLMs Fail at IaC
1. The Versioning Puzzle
Cloud providers change APIs faster than LLMs retrain. An AI trained on 2024 data will generate `acl="private"` for S3 buckets, unaware that AWS deprecated ACLs. The result? Immediate deployment failure.
2. State Drift Blindness
LLMs treat code as the source of truth. But in IaC, the *real* source of truth is the live cloud state. Generic AI suggests "just delete the resource" code blocks, oblivious to the production database attached to it.
3. Permission Hell
To avoid errors, AI defaults to `*:*` permissions. It's easier to give AdminAccess than to figure out the exact `kms:Decrypt` action needed. This creates a security nightmare masked as "working code."
4. Secret Leakage
"Here, put your API key in this variable." No. Stop. The AI doesn't understand KMS, Vault, or Pulumi ESC unless explicitly forced to respecting secret providers.
The Experiment: Fewer Failed Redeploy Cycles
We reverse-engineered that "Dashboard AI" competence into a reusable SKILL.md for Google Antigravity. Then we ran a controlled test with 4 groups of entry-level devops engineers (because let's be real, seniors don't have time for this).
- Groups 1 & 2Standard Cursor Rules. Immediate chaos. Struggles with state locks and circular dependencies.
- Group 3Armed with the Antigravity Pulumi Skill. Detected drift autonomously.
- Group 4Control group (No AI). Pure manual suffering.
The result? Group 3 spent substantially less time waiting for `pulumi up` to fail than the rules-only groups in the same exercise window.
The "Intern Study"
Average redeploys required to fix a drifting Pulumi stack.
The Pulumi Mastery Skill
This isn't just a prompt. It's a procedural standard for managing Python-based Pulumi architecture. It handles programmatic drift detection (`stack.refresh()`), enforces KMS usage, and validates provider versions.
If you ship Pulumi from AI-assisted editors, scope the model with refresh/cancel flows, KMS defaults, and provider version checks. Generic prompts will happily suggest *:* IAM and deprecated S3 ACL arguments.