Software Engineer 3, Platform
Welcome to Our World—we’ve been leading the charge in the affiliate industry from day one, establishing performance marketing and paving the way for future innovations. We're known for maintaining one of the largest, most reliable partnership platforms with impeccable, personalized service. Founded in Santa Barbara, California in 1998, CJ (formerly Commission Junction) stands as the most trusted name in performance marketing. We specialize in building partnerships between top brands and reputable publishers to drive revenue and business growth. CJ’s industry-leading solutions make us the platform of choice for over 3,800 global brands across sectors like retail, travel, finance, technology, and home services. As part of Publicis Groupe, our savvy data capabilities, cutting-edge tech, and strategic expertise facilitate genuine connections, allowing brands to reach consumers wherever they are.
A Quick Peek at Affiliate Marketing: Think back to your last online purchase. Did an influencer tip you off about a great product and offer a discount? Or perhaps you relied on a trusted review site to make your decision? Whatever path you took, affiliate publishers likely played a role by influencing, informing, or helping you find the best deal. CJ connects brands with these publishers, creating valuable resources for shoppers like you.
You must be work authorized in the United States without the need for employer sponsorship. This is a hybrid role requiring 3 days a week in office.
About CJ Engineering
At CJ, we are passionate about software engineering. We build exciting software with quality and maintainability in mind, believing in common sense, simplicity, and efficiency. We practice critical thinking, challenge each other regardless of title, and value the wisdom of the team—that makes us Engineers, not just developers. Here are some principles that set CJ engineers apart:
- Engineering autonomy: Business decisions are made by business people, and technical decisions are made by technical people.
- Full stack: Expect to gain competence in every aspect of software engineering, from frontend to database, requirements analysis, testing, and technology selection.
- Clean, maintainable code: Code is read more often than it is written, so we prioritize writing it well from the start.
- Pairing: The highest quality code comes from close collaboration, so we pair by default.
- TDD: Quality is baked into our process through Test-Driven Development.
- Ownership: Engineers own the full lifecycle of what they build—from design and implementation to deployment, monitoring, production support, and on-call rotations.
- Operational Excellence: We embrace Infrastructure as Code, CI/CD, automation, and observability to build reliable systems and deliver software safely, efficiently, and at scale.
We believe in Agile values and incremental development, constantly experimenting, retrospecting, and adjusting. We are committed to exploring how AI can amplify our productivity, viewing it as a force multiplier, not a replacement for good engineering.
As a Software Engineer 3 on the Engineering Experience (EngExp) platform team, you help run and evolve the platform powering CJ's production systems across multiple AWS regions. "Platform" here is broad—it includes Kubernetes clusters, the observability stack every squad depends on, CI/CD and artifact infrastructure, AWS networking, secrets and access systems, and cost visibility. EngExp owns all of it. This is not just an infrastructure role; your value lies in engineering judgment: evaluating systems, detecting risk, and making decisions under uncertainty. You'll own meaningful pieces of these systems independently and drive changes from design through production.
Responsibilities
What systems you will work on:
- Observability & monitoring: Prometheus, Alertmanager, Grafana, and OpenTelemetry across production regions. This is not just dashboard-building—you'll own cardinality budgets, recording-rule design, keep Prometheus healthy as it scales (federation/sharding/long-term storage), and understand Alertmanager HA and alert-routing config. Deep Prometheus and Alertmanager knowledge is a core requirement.
- Kubernetes & cloud infrastructure: Multi-region EKS clusters, including upgrades, node group and Karpenter management, controller lifecycle, and add-on/configuration management. Spot failure modes before they happen (e.g., subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption).
- AWS networking: VPC and subnet design, CIDR management, VPC peering, Route53, security groups, and NAT gateway topology across accounts and regions, plus 24/7 networking alarms for production traffic between clusters and squad resources.
- CI/CD & artifact management: GitLab administration (runner fleet, cache, access—not just pipeline authoring), GitOps delivery through ArgoCD, and Nexus artifact repository, including storage lifecycle management.
- Access & identity: Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management. Turn recurring access requests into self-service workflows that are hard to misuse.
- Cost observability: OpenCost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution across teams.
What You'll Do:
- Own and evolve meaningful pieces of the systems above, focusing on what is happening and why.
- Manage infrastructure-as-code with Terraform across AWS accounts.
- Build and maintain GitLab CI/CD pipelines and GitOps delivery (ArgoCD).
- Enforce platform standards: RBAC, admission webhooks, resource limits, LimitRanges.
- Turn recurring requests (ingress, DNS, service accounts) into self-service workflows.
- Respond to and help drive resolution of platform incidents, focusing on learning and system improvement.
- Act as a reviewer of infrastructure changes—Terraform, Kubernetes configs, observability config.
Technologies We Use
- Kubernetes / EKS (multi-cluster, multi-region), Karpenter, cert-manager, external-dns
- Prometheus, Alertmanager, Grafana, OpenTelemetry (and long-term storage/sharding for Prometheus)
- AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions)
- Terraform, AWS (IAM, EKS, S3, EBS)
- ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines
- Vault, OpenCost
- Kubernetes controllers/operators (reconciliation patterns, restart safety)—Go experience is a plus, not required
Qualifications
What We Look For:
- 3+ years of experience in software or infrastructure engineering.
- Bachelor's degree or equivalent experience.
- Hands-on production experience with Kubernetes and at least one major cloud (AWS preferred).
- Real operational depth in at least one system we own beyond the cluster—most valuably the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We filter for people who have run these systems, not just used them.
- Comfortable owning infrastructure-as-code (Terraform) and CI/CD pipelines.
- Ability to reason about tradeoffs and communicate the pros and cons of multiple approaches.
- Effective communication; thrives in a collaborative, pair-friendly team culture.
Nice To Have:
- AWS networking depth (Transit Gateway, multi-account topology).
- Prometheus long-term storage/sharding (Thanos, Cortex, Mimir, or equivalent).
- Kubernetes controllers/operators—Go experience is a plus, not required.
- Policy-as-code (Kyverno/OPA).
What Success Looks Like
- You own platform components independently and ship changes that are intentional and low-risk.
- The systems you own are understood deeply enough that we stop making decisions we have to reverse.
- Production issues are understood quickly because of the observability and instincts you bring.
- Engineers can deploy and debug services with less platform intervention over time.
Benefits
- Competitive salaries and 401K matching.
- Comprehensive medical, dental, and vision coverage.
- Flexible time off without accrual hassles.
- A generous number of paid holidays.
- Company-sponsored team-building events.
- Employee Referral Program.
- Annual recognition awards.
- Hybrid work arrangements for optimal work-life balance.
- Parental bonding leave.
- Backup care options for children and elders.
- Employee discount program.
- International SOS program for global support.
- Business Resource Groups, where employees connect over shared interests to cultivate an engaging, inclusive environment.
Pay
Compensation Range: USD $88,540.00 - USD $135,632.00/Annually. This is the pay range the Company believes it will pay for this position at the time of posting. Compensation will be determined based on skills, qualifications, and experience, along with the requirements of the position.
Schedule
This is a hybrid role requiring 3 days a week in office.