Staff Software Engineer - Backend & AI Infra
MLabs · United States · 4 mo ago
RemoteRemoteInformation TechnologyFull-time
Key Responsibilities
- Plugin Runtime Ownership: Lead the evolution of the per-agent process, migrating from a distributed Go/Python hybrid to a centralized, high-performance Go service utilizing Postgres state and real-time websocket price feeds.
- Rules Engine Engineering: Build a YAML-configurable "Scanner Gateway" to bridge signal production and execution, allowing for complex scoring and filtering without direct code manipulation.
- Advanced Execution Systems: Develop and maintain the RatchetStop Backend, a centralized profit-trailing service capable of sub-second evaluation and websocket-based order execution to protect capital even when agents are offline.
- Data & Connectivity: Manage the Model Context Protocol (MCP) server bridging agents to platform tools, and oversee a high-throughput data pipeline (Redis, Postgres, ClickHouse) for real-time market intelligence ingestion.
- Model & Agent Hosting Migration Infrastructure Sovereignty: Lead the technical execution of migrating agents from third-party platforms to a custom-built, Senpi-hosted environment featuring isolated workspaces and state persistence.
- Model Serving: Evaluate and implement the transition from external LLM APIs (Anthropic, Google) to self-hosted inference, optimizing for telemetry capture and performance.
- Telemetry & Feedback Loops: Architect systems to capture every agent decision and score, creating a self-reinforcing loop where the fleet learns and improves from collective performance data.
- Deployment Pipelines: Build robust CI/CD pipelines for zero-downtime rollouts, ensuring that updates to scanner logic or runtime patches do not interrupt active market positions.
- Infrastructure & Operations System Reliability: Design monitoring and alerting frameworks to detect agent failures, state corruption, or authentication expirations before they impact financial performance.
- Cloud Orchestration: Manage AWS/EKS environments using Infrastructure-as-Code (IaC).
- Incident Response: Own the operational health of the fleet, acting as the primary responder for high-stakes trading system incidents.
Requirements
- Technical Essentials: Expert Backend Engineering, Startup Experience, Real-Time Systems, Database Mastery, End-to-End Ownership.
- Preferred Qualifications: LLM Infrastructure, FinTech/Trading, Agentic Frameworks.
Benefits
- Competitive compensation and equity packages.
- The opportunity to build foundational infrastructure in a new category of autonomous software.
- A high-autonomy environment with a focus on engineering excellence.
- A collaborative culture working alongside industry-leading founders and engineers.