Senior QA Engineer
About the role
We're looking for a Senior QA Engineer to own the health of every release. You own the release gate: if your suite is green, we ship; if it's red, we don't. Nobody overrides a red gate by asserting that it's probably fine; they either fix the product or fix the test with you. Your real product is a signal the whole company believes without checking. You'll build that signal with AI as your primary lever, not as a garnish.
We're a small, distributed, AI-first team — this is a high-autonomy, high-consequence role, not a ticket queue. Our stack: Python/Django + DRF backend on Postgres and Redis, a TypeScript/Next.js customer-facing app, GitLab CI, Terraform, Docker, and third-party carrier, airline, and tracking integrations that misbehave in unpredictable ways.
Responsibilities
- The release gate: Define what must be true before code reaches customers, automate that decision, and keep runtime short enough that nobody wants to skip it.
- API test automation: Broad, fast, deterministic coverage of our Django/DRF surface — contracts, auth and permission boundaries, pagination, error shapes, idempotency, and integration seams with third-party APIs.
- End-to-end front-end automation: Real browser coverage of revenue-generating and support-generating journeys, using a modern framework like Playwright, stable enough to run on every merge.
- AI-driven test generation and maintenance: Build pipelines that turn specs into first-pass suites, keep suites current, and draft fixes for stale selectors or fixtures. Automate maintenance to reduce cost.
- Failure triage: Ensure the team receives actionable failure reports — the failing assertion, likely responsible diff/deploy, and a first hypothesis — not just a log link. Increasingly, this triage starts with an agent.
- Flake as a defect class: Track flake rate as a real metric with a real budget. Quarantine, root-cause, and fix — never silence. Muted tests must expire.
- Environments and test data: Maintain reproducible environments and seeded, realistic fixtures so failures are meaningful and passes aren’t luck.
- Non-functional coverage: Performance regressions on critical endpoints and accessibility checks in the front-end suite.
- Production verification: Smoke and synthetic checks after every release, with a closed loop from escaped bugs back into the suite.
- Quality as a shared practice: Own the gate, but raise the floor for the whole team — review testability in specs, make tools easy to use, and coach engineers and agents to leave coverage behind.
What success looks like
- First 30 days: Understand the release process, current coverage, and existing failures. Replace the riskiest manual pre-release check with an automated one and ship it.
- First 60 days: A single gate command runs on every release with meaningful API and end-to-end coverage, in a runtime nobody argues with. Flake rate is measured and visible. AI-assisted generation produces real tests, not demos.
- First 90 days: The gate is trusted enough that a red result stops a release without discussion. Escaped defects trend down, each has a test, and new work arrives with coverage.
Qualifications
- 5+ years in QA, SDET, or test-automation engineering, including ownership of an automated suite a team genuinely depended on to release.
- Strong hands-on coding in Python and/or TypeScript — enough to read application code, not just drive it from outside.
- Deep API testing experience against REST services, ideally Django/DRF: auth, permissions, contract validation, and integration seams with unreliable third parties.
- Real end-to-end browser automation experience (Playwright, Cypress, or equivalent) with a track record of keeping suites stable.
- CI/CD fluency: Building and owning pipeline stages, parallelization, artifacts, reporting, and gating releases. GitLab CI is a plus.
- Demonstrated daily use of AI tools in your workflow, with concrete examples — ideally including test generation or maintenance.
- Comfort with containers and cloud environments (Docker, Terraform) to stand up and debug environments yourself.
- Excellent written English.
Nice to have: performance/load testing, accessibility testing, contract testing, synthetic monitoring (we use Datadog), or a security-testing background.
Skills
- Accountability: You want ownership of the release gate and are comfortable saying "not yet" with evidence.
- Automation-first mindset: You treat test code as production code — reviewed, refactored, and held to the same bar.
- Skeptical by instinct: You ask what a green result actually proved and prefer to break your own checks to verify they work.
- AI-fluent: You already use LLMs and coding agents daily, know their strengths/weaknesses, and have a verification habit for AI-generated tests.
- Excellent communicator: You write clearly for engineers and stakeholders, and failure reports tell readers what to do next.
- Systems thinker: You fix the class of failure, not the instance, and see bugs as signals of missing coverage.
- Pragmatic about coverage: You focus on risk, not percentages.
- Independent and async-native: Comfortable working across time zones without supervision, proactive about context, and reliable with commitments.
- Low-ego and hands-on: You’ll write fixtures, chase flakes, fix CI jobs, and pair with engineers.
- Balanced: You work hard during the day and value personal time, stepping in only for genuine urgency.
Benefits
- Genuine authority over what ships in a company that wants to be told "no" when it’s the right answer.
- Build a quality function from a clean slate with modern tooling — no legacy Selenium or untrusted suites.
- Work in an environment where AI tooling is genuinely embraced, not tolerated.
- Real ownership and short decision paths: no QA-vs-engineering politics, no committees.
- Visible impact: your work directly affects shipments in the physical world.
- Flexible, async-friendly, globally distributed team that respects personal time and values results over hours.