Infrastructure Engineer
Anthelion Capital · New York, NY · 3 mo ago
HybridInformation Technology$140k–$225k/yrFull-time
About the role
Anthelion is a next-generation investment firm building a proprietary AI and data platform that powers our investment lifecycle from underwriting to portfolio management. The platform integrates structured and unstructured data, advanced analytics, and automated workflows to drive superior, risk-adjusted returns in private credit and structured finance. We are engineers and investors working together to redefine how institutional investment decisions are made faster, smarter, and more transparent.
Responsibilities
- Deploy, configure, and maintain shared platform services as containerized workloads including end-to-end ownership of networking, access, and connectivity between services.
- Manage cloud infrastructure, including container registries, managed identities, Key Vault secrets, storage backends, and virtual network configurations.
- Build and maintain CI/CD pipelines, branch protection policies, and release management workflows across repositories.
- Continuously evaluate and adopt tools and technologies that improve platform reliability, developer experience, and team velocity.
- For those interested in growing into AI systems work, there is real room to do so over time — though none of the following is required to be successful in this role:
- Support the buildout and operationalization of agentic AI workflows, including agent hosting, lifecycle management, and integration with Model Context Protocol (MCP) servers.
- Help build shared tooling and infrastructure that enables data scientists to develop, test, and deploy agents with minimal friction.
- Contribute to evaluation frameworks and quality standards for AI agents, including automated benchmarking, regression testing, and production-readiness criteria.
- Extend observability and reliability practices into agent execution environments, including logging, tracing, and performance monitoring.
Requirements
- 3+ years of experience in infrastructure, DevOps, platform engineering, or SRE roles with a clear track record of building and maintaining production systems.
- Solid understanding of containerization and cloud infrastructure — Docker, Kubernetes, and at least one major cloud provider.
- Hands-on experience deploying and operating containerized services in cloud environments, including configuring networking, load balancing, and service-to-service connectivity.
- Experience building and maintaining CI/CD pipelines, Git-based release management, and branch protection workflows.
- Experience with workflow orchestration tools (Prefect, Airflow, Dagster, or similar) in production environments.
- Familiarity with monitoring and observability tooling health metrics, alerting, logging, and tracing.
- Strong documentation habits and the ability to communicate technical architecture clearly to diverse stakeholders.
Qualifications
- A genuine interest in AI and a desire to learn and grow into building, hosting, and operating AI agents and agentic systems.
- Familiar with agentic workflow frameworks (e.g., MCP, LangChain, or similar).
- Experience with MLOps or ML infrastructure, including model training, retraining, and inference workflows.
- Familiarity with model serving and deployment patterns (batch inference, real-time APIs, feature stores).
- Experience standing up and maintaining third-party AI/ML platform tools (e.g., Langfuse, MLflow, or similar observability and evaluation platforms).
- Experience managing internal Python package distribution (private PyPI, Artifactory, or similar).
- Openness to flexing into adjacent engineering work data engineering, software engineering, and similar to help fill extra capacity where the team needs it.
Benefits
- Comprehensive health, dental, and vision insurance.
- Retirement savings plan with company match.
- Hybrid/flexible work arrangements and a supportive work environment.
Pay
$140,000 to $225,000 per year
Schedule
Hybrid/flexible work arrangements