You will take over a working platform, not build one. Rabbit already provisions and runs production for a dozen clients. Your job is to keep it current, extend it when a client needs something new, and make every change land as a pull request that another engineer could read and revert.
The work is concrete: a base image needs a patch and every tenant that runs on it needs a redeploy; a client needs a new managed service wired into their environment; a monitoring threshold is wrong across eighteen repositories; a migration that was started needs to be finished and rolled out. You will be handed a GitHub issue with enough context to start, and you are expected to leave commits, comments and a paper trail behind.
This is a contract role, remote, with overlap required for US Eastern working hours. You will hold administrative access to client infrastructure, so we hire slowly and check references.
UDX is a service-disabled veteran-owned small business (SDVOSB) founded in 2011 and headquartered in Durham, North Carolina, with a distributed engineering team. We build and run cloud infrastructure and web platforms for organizations that cannot afford downtime: cultural institutions, live event promoters, consumer brands, real estate and professional certification bodies. Our delivery platform, Rabbit, turns a Git repository into a fully provisioned, monitored, secured environment on Kubernetes across GCP, AWS and Azure. We also ship Kubernetes operators to the Azure Marketplace.
Every hour we work maps to a GitHub issue. Every change is a pull request. Every day of work leaves a comment. That is how we bill, how we prove work to clients, and how we know what happened six months later.
What you'll do
- Own the worker image family (PHP, Node.js, WordPress runtime, SFTP gateway): base image updates, security patches, release tagging, and the automated redeploy of every tenant running on them.
- Maintain and extend the Rabbit automation modules (OpenTofu/Terraform on GCP, AWS and Azure) and the reusable GitHub Actions workflows that every client repository consumes.
- Finish and roll out platform migrations across tenant repositories, one pull request per repository, with the rollout tracked in a single issue.
- Run cloud security operations for clients: Security Command Center and GuardDuty findings, IAM and workload identity, secrets inventory, monthly OS patching, backup and restore tests.
- Keep the dependency and image update bots honest: verify what they change, catch what they miss, and never let a green dashboard stand in for a merged pull request.
- Diagnose production incidents on GKE and AKS: pod scheduling, cost spikes, DNS and certificate expiry, cache invalidation, database connectivity.
- Write the operational record as you go: an issue comment per working day, a PR per change, a runbook per recurring task.
- Five or more years running production Kubernetes, with at least two years on GKE, EKS or AKS in a multi-tenant or multi-client setting.
- Fluency in OpenTofu or Terraform modules, GitHub Actions (reusable workflows, composite actions, OIDC federation) and Docker image supply chain.
- Working knowledge of GCP and AWS IAM, workload identity federation, and cloud-native secret management. Azure is a plus.
- Enough PHP and Node.js runtime knowledge to debug why a WordPress or Next.js container is unhealthy.
- A habit of leaving evidence: you link issues in commits, comment before you close, and can explain any change from the log alone.
- Written English that a client can read. You will be in their repositories.
- Overlap with US Eastern hours for calls and incidents.
- Contract engagement, hourly, with a defined scope and a real backlog on day one.
- A platform that is already built and running in production, so your work ships to real clients immediately.
- Multi-cloud work (GCP, AWS, Azure) and Kubernetes operator development destined for the Azure Marketplace.
- Direct access to the founders. No layers, no ticket queue between you and the decision.
- Remote, flexible hours outside the US Eastern overlap.
| Specialization | Cloud platform, Kubernetes, infrastructure as code, security operations |
|---|---|
| Commitment | 20 to 30 hours a week to start, growing with the backlog |
| Engagement | Contract, hourly, invoiced monthly against GitHub issues |
| Work setup | Remote |
| Overlap | At least 4 hours a day inside US Eastern business hours |
| Access | Administrative access to client infrastructure, granted after reference checks, hardware key required |
| Location | No restriction; US or European time zones preferred for incident overlap |
| Language | English, written and spoken; you will write in client repositories |
| Start | As soon as the references clear |
| How we work | Issue per task, pull request per change, comment per day |
Send the answers below through the contact form with the role name in the subject line. Short answers are fine; links are better than descriptions. We reply to everyone who answers all four.
- Link to a pull request you are proud of on a Kubernetes or infrastructure-as-code codebase, and say what it fixed.
- Which of the responsibilities above have you done in production, for whom, and for how long?
- Describe an incident you diagnosed on GKE, EKS or AKS: what broke, how you found it, what you changed afterwards.
- Your availability in hours per week, the US Eastern hours you can overlap, and your hourly rate.