Senior Software Engineer, Platform
Welcome to Our World — we’ve been leading the charge in the affiliate industry since 1998. As CJ (formerly Commission Junction), we’re the most trusted name in performance marketing, specializing in building partnerships between top brands and reputable publishers to drive revenue and business growth. Our industry-leading platform serves over 3,800 global brands across retail, travel, finance, technology, and home services. Backed by Publicis Groupe, we leverage savvy data capabilities, cutting-edge tech, and strategic expertise to facilitate genuine connections and help brands reach consumers wherever they are.
Think of affiliate marketing as the behind-the-scenes force influencing your online purchases—whether through influencers, review sites, or discount offers. CJ connects brands with these publishers, creating valuable resources for shoppers like you.
You must be work authorized in the United States without the need for employer sponsorship. This is a hybrid role requiring 3 days a week in a CJ office location.
About CJ Engineering
At CJ, we are passionate about software engineering. We build exciting software with a focus on quality and maintainability, valuing common sense, simplicity, and efficiency. Our engineers practice critical thinking, challenge each other regardless of title, and believe in the wisdom of the team. Here’s what sets CJ engineers apart:
- Engineering autonomy: Business decisions are made by business people, and technical decisions are made by technical people.
- Full stack: Gain competence in every aspect of software engineering—frontend, database, requirements analysis, testing, and technology selection.
- Clean, maintainable code: Code is read more often than it’s written, so we prioritize writing it well from the start.
- Pairing: High-quality code comes from close collaboration, so we pair by default.
- TDD: Quality is baked into our process through Test-Driven Development.
- Ownership: Engineers own the full lifecycle of what they build—from design and implementation to deployment, monitoring, production support, and on-call rotations.
- Operational Excellence: We embrace Infrastructure as Code, CI/CD, automation, and observability to build reliable systems and deliver software safely, efficiently, and at scale.
We believe in Agile values and incremental development, constantly experimenting, retrospecting, and adjusting. We’re committed to exploring how AI can amplify our productivity, viewing it as a force multiplier—not a replacement for good engineering.
About the Role
As a Senior Software Engineer on the Engineering Experience (EngExp) platform team, you’ll help run and evolve the platform powering CJ’s production systems across multiple AWS regions. The “platform” is broad—it includes Kubernetes clusters, the observability stack, CI/CD and artifact infrastructure, AWS networking, secrets and access systems, and cost visibility. EngExp owns all of it, and this role touches most of it.
This is not just an infrastructure role—your value lies in engineering judgment. You’ll act as a leader on the team, owning the reliability and operability of the platform, setting standards, and mentoring other engineers. We’re looking for someone with deep expertise in the systems we own, as we’ve had to reverse platform decisions in the past due to a lack of depth in critical components.
Responsibilities
The EngExp team owns the following systems, and you’ll be responsible for their reliability, critical review of changes, and raising the team’s depth in them:
- Observability & monitoring: Prometheus, Alertmanager, Grafana, and OpenTelemetry across every production region. This goes beyond dashboard-building—you’ll own cardinality budgets, recording-rule design, Prometheus health (including federation/sharding/long-term storage strategies), and Alertmanager HA to minimize blast radius.
- Kubernetes & cloud infrastructure: Multi-region EKS clusters, including upgrades, node group and Karpenter management, controller lifecycle, and add-on/configuration management. You’ll identify failure modes (e.g., subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption) before they happen.
- AWS networking: VPC design, subnet allocation and CIDR management, VPC peering, Transit Gateway, security groups, Route53, and NAT gateway topology across multiple accounts and regions, plus 24/7 networking alarms for all production traffic.
- CI/CD & artifact management: GitLab administration (runner fleet, AMI updates, cache, access), GitOps delivery via ArgoCD, and Nexus artifact repository (including storage lifecycle management).
- Access & identity: Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management (including cost alerts and reporting).
- Cost observability: OpenCost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution across teams to ensure waste is attributable rather than shared overhead.
- Internal tools & delivery: Container image build pipeline and base image standards, code audit tooling, HedgeDoc, UI CDN (S3 + CloudFront), and adopted applications with no other owner.
What you’ll do:
- Own the reliability and operability of the systems above, focusing on what is happening and why.
- Establish and enforce platform standards: RBAC, admission webhooks, resource limits, LimitRanges, and policy-as-code.
- Manage infrastructure-as-code with Terraform across AWS accounts.
- Act as a high-quality reviewer of infrastructure changes—Terraform, Kubernetes configs, CI/CD pipelines, and observability config—to catch subtle issues and long-term risks before they ship.
- Turn recurring requests (ingress, DNS, service accounts) into self-service workflows that are hard to misuse.
- Drive resolution of platform incidents with a focus on learning and lasting system improvement.
- Evaluate new patterns (e.g., Gateway API/Kgateway, claim-based self-service) based on tradeoffs, not hype.
- Mentor less-senior engineers and raise the team’s depth in the components we own.
Technologies We Use
- Kubernetes / EKS (multi-cluster, multi-region), Karpenter, cert-manager, external-dns
- Prometheus, Alertmanager, Grafana, OpenTelemetry (and long-term storage/sharding for Prometheus)
- AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions)
- Terraform, AWS (IAM, EKS, S3, EBS)
- ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines
- Vault, OpenCost
- Gateway API / Kgateway
- Kubernetes controllers/operators (reconciliation patterns, restart safety)—Go experience is a plus, not required
Qualifications
- 6+ years of experience in software and/or infrastructure engineering.
- Bachelor’s degree or equivalent experience.
- Deep, hands-on production experience operating Kubernetes and AWS at scale, across multiple accounts and regions.
- Real operational depth in at least one system we own beyond the cluster—most critically the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We’re filtering for people who have run these systems, not just used them.
- Strong AWS networking judgment (VPC, peering, Transit Gateway, subnet/CIDR design).
- A track record as a critical reviewer—spotting subtle infrastructure issues and long-term risks before they ship.
- Experience leading technical work and mentoring engineers; ability to manage, clarify, and plan around uncertainty.
- Effective communication and the ability to influence design in a product-focused way.
Nice to have:
- Prometheus long-term storage/sharding (Thanos, Cortex, Mimir, or equivalent) run in production.
- Experience owning a container image/base image pipeline.
- Policy-as-code (Kyverno/OPA) and admission webhook design.
- Building claim-based self-service platform capabilities.
What Success Looks Like
- Engineers can deploy and debug services without needing platform intervention.
- The systems we own are understood deeply enough that we stop making decisions we have to reverse.
- Production issues are understood quickly, not just reacted to.
- Platform changes are intentional and low-risk, and standards are clear and enforced.
Benefits
- Competitive salaries and 401K matching.
- Comprehensive medical, dental, and vision coverage.
- Flexible time off without accrual hassles.
- Generous number of paid holidays.
- Company-sponsored team-building events.
- Employee Referral Program.
- Annual recognition awards.
- Hybrid work arrangements for optimal work-life balance.
- Parental bonding leave.
- Backup care options for children and elders.
- Employee discount program.
- International SOS program for global support.
- Business Resource Groups for employees to connect over shared interests and cultivate an engaging, inclusive environment.
Pay
Compensation range: USD $112,290.00 – USD $172,032.00/year. This range reflects the pay the Company believes it will pay for this position at the time of posting. Compensation will be determined based on skills, qualifications, and experience, along with the requirements of the position. The Company reserves the right to modify this range at any time.
Schedule
This is a hybrid role requiring 3 days a week in a CJ office location.