Analyst, Application Engineer
Join one of the top FinTech companies, where our Aladdin platform serves over 200 leading financial institutions managing approximately a quarter of global assets under management. BlackRock is a global, close-knit organization united by a mission to deliver exceptional service to our business partners and customers. We value diversity of thought, background, and experience, and invest in our people through flexible time off, collaborative work environments, and strong career development opportunities.
About the role
In this role, you will support business-critical computing workloads, including real-time and batch processing, data transfer services, application onboarding and upgrades, and recovery procedures. You will be part of a globally distributed team operating 24x7x365 to ensure the stability and reliability of production environments. This position offers hands-on exposure to large-scale production systems and modern operational practices, such as automation, observability, and AI-assisted operations.
Team Overview
The Service Management Operations Group monitors, supports, and administers production environments for all BlackRock businesses, including subsidiaries and BlackRock Solutions. The group acts as a first responder for incident detection, troubleshooting, resolution, and escalation. You will collaborate with experienced professionals across regions and technologies, gaining broad exposure to platforms and applications while contributing to service quality, reliability, and continuous operational improvement as part of the One BlackRock culture.
Responsibilities
- Support reliability and availability of production systems by monitoring environments and responding to alerts and incidents according to documented procedures.
- Assist in maintaining availability, performance, and recovery objectives for critical workloads.
- Participate in incident reviews and help document root causes and follow-up actions.
- Use monitoring and observability tools to identify system health issues and potential risks, validating alerts and distinguishing real production impact from noise.
- Escalate issues appropriately based on impact, urgency, and runbooks.
- Execute scripted remediation and automation for known failure scenarios while following defined guardrails, approvals, and audit requirements.
- Identify recurring manual tasks or failure patterns and suggest candidates for automation.
- Work with engineering teams during deployments, upgrades, and production readiness activities, providing operational feedback on monitoring gaps, documentation quality, and supportability issues.
- Help ensure services are observable, recoverable, and supportable in production.
- Support change implementation by monitoring post-change system behavior and health.
- Assist with capacity, resilience, and disaster recovery activities such as testing and exercises.
- Follow disciplined change and incident management processes.
- Maintain accurate incident records, handover notes, and operational documentation.
- Support audit and compliance requests by gathering logs, metrics, and operational evidence.
- Contribute to continuous improvement through post-incident learnings and process updates.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
- 0–3 years of experience in production support, service management, operations, DevOps, or related technical roles.
- Basic familiarity with Linux/Unix systems, networking concepts, and distributed applications.
- Exposure to monitoring/observability tools (metrics, logs, dashboards, alerts).
- Understanding of incident management fundamentals and IT service management concepts.
- Interest in automation, scripting, and modern operational practices (e.g., reliability, resiliency).
- Strong analytical skills, attention to detail, and ability to follow structured processes.
- Clear written and verbal communication skills; comfortable working in a global, follow-the-sun team.
- Willingness to learn, adapt, and operate in a 24x7 production support environment.
Benefits
- Strong retirement plan.
- Tuition reimbursement.
- Comprehensive healthcare.
- Support for working parents.
- Flexible Time Off (FTO) to relax, recharge, and spend time with loved ones.
Schedule
Employees are currently required to work at least 4 days in the office per week, with the flexibility to work from home 1 day a week. Some business groups may require more time in the office due to their roles and responsibilities. New joiners can expect this hybrid model to accelerate learning and onboarding.