Production Reliability Engineer
About the Company
Friedman Vartolo LLP is a fast-growing, New York-based law firm specializing in real estate and default services, with over 300 employees providing top-tier legal services to clients in seven states. While our legal expertise sets us apart, it's our mindset that drives us forward. We bring a fresh, fast-paced energy that shapes how we approach every challenge. We solve problems at the root rather than settling for surface fixes. Every individual adds value, owns outcomes, and moves the firm forward. With an underdog mentality, we embrace constant elevation, always sharpening and climbing. When challenges arise, we collaborate as one team to get the job done.
About the Role
We are seeking a Production Reliability Engineer to own the operational health of our deployed internal application portfolio. The portfolio is growing rapidly—many new applications ship every week—and we need a dedicated owner ensuring deployed applications stay reliable, secure, and performant over their lifetime while keeping the operational footprint manageable. This is a hands-on role for an engineer who thinks systematically about reliability, communicates clearly during incidents, and knows when to retire applications that have outlived their usefulness. You will be the source of truth for "is our portfolio healthy" and the decision-maker on what gets maintained, fixed, or retired.
You will work closely with the Senior Technology Manager, a DevOps Engineer (counterpart role), and a group of Production Engineers who own initial deployment and 30-day iteration of each application.
Responsibilities
- Own the operational health of the firm's deployed internal application portfolio
- Lead incident response for production issues; serve as on-call rotation lead
- Manage security patching, dependency updates, base image refreshes, and certificate rotation across the deployed portfolio
- Conduct monthly triage of every deployed application—usage levels, error rates, security posture, retire/keep recommendation
- Maintain the application retirement queue and drive retire/deprecate decisions for low-usage or obsolete applications
- Define and track Service Level Objectives (SLOs) for application uptime, performance, and reliability
- Partner with the DevOps Engineer to ensure deployment pipelines incorporate reliability requirements from day one
- Produce a weekly portfolio health report—uptime trends, open incidents, security posture, retire/deprecate decisions
- Conduct incident postmortems and drive remediation actions to completion
- Contribute to security and compliance posture work (SOC2 readiness, Vanta evidence collection) as it relates to operational reliability
Requirements
- 5+ years professional experience in Site Reliability Engineering, DevOps, production engineering, or equivalent
- Strong incident-response experience—leading incidents, not just participating; comfortable owning the on-call rotation
- Experience defining and tracking SLOs and error budgets for production systems
- Strong Azure cloud experience (or directly comparable AWS/GCP experience)
- Hands-on experience with monitoring, alerting, and observability tooling (Azure Monitor, Application Insights, Datadog, New Relic, or equivalent)
- Comfortable reading code in multiple languages (TypeScript, Python, C# preferred) for incident triage and root-cause analysis
- Understanding of cloud security fundamentals—patching cadence, dependency management, secret rotation, base image hygiene
- Excellent written communication for incident comms, postmortems, and portfolio reporting
Preferred Qualifications
- Experience managing a portfolio of small-to-medium applications, rather than a single monolith
- Experience with security tooling (Huntress, Microsoft Defender, Vanta, similar)
- Familiarity with Databricks, data warehouses, or analytics platforms
- Experience driving retire/deprecate decisions and the accompanying conversations
- Experience in regulated industries (legal, financial services, healthcare)
- Familiarity with compliance frameworks (SOC2, ISO 27001) from an operational-evidence perspective
Benefits
- Paid parental leave options
- Short and long-term disability leave options
- Comprehensive medical, dental, and vision insurance
- Flexible Spending Accounts (FSA) and Dependent Care plans
- Commuter benefits including transit and parking options
- Pet insurance
- 401(k) retirement plans with employer contributions
- Gym and fitness reimbursements
- Annual business expense reimbursements for attorneys
- Annual $500 travel budget to attend firm-sponsored social events
- Hybrid work flexibility available after 90 days of employment (dependent upon specific position requirements and firm business needs)