Site Reliability Engineer – Investment FundOur client is a specialized investment firm in New York looking to hire a Site Reliability Engineer as the second engineer on its technology team. The firm is a lean, approximately 17-person investment platform focused on micro-cap PIPE and structured equity investments. Its technology stack has been built internally and includes AI applications used daily by the investment team, an internal knowledge system, scheduled data pipelines, and infrastructure supporting the firm's trading systems. This hire will work directly with the CTO and take ownership of production reliability across the platform. This is a broad, hands-on role with no layers between engineering and the investment team.
What You'll Do:
  • Own day-to-day reliability of production applications, data pipelines, AI tools, and trading-related systems
  • Build and maintain monitoring, alerting, logging, and tracing across the technology environment
  • Monitor scheduled data pipelines and troubleshoot late or missing data, failed dependencies, retries, and reruns
  • Serve as the first line of technical support for the firm's trading systems
  • Diagnose and resolve production incidents across applications, infrastructure, data, and integrations
  • Read, debug, and patch production Python code when issues arise
  • Support Linux-based applications, shell scripts, cron jobs, and scheduled processes
  • Operate applications and data workflows across Azure and AWSSupport deployment, backups, database connections, performance issues, and other day-to-day cloud operations
  • Monitor and troubleshoot LLM-backed applications, including provider outages, rate limits, model changes, silent agent failures, and unexpected token usage
  • Automate recurring operational work and eliminate manual processes
  • Build runbooks and documentation for systems currently maintained by a small engineering team
  • Communicate directly with investment, trading, operations, and senior leadership when issues arise
  • What We're Looking For4–8 years of experience in Site Reliability Engineering, Production Support, Application Support, Infrastructure Engineering, Platform Engineering, or a similar hands-on production role
  • Strong Linux, shell scripting, and cron experience
  • Strong working knowledge of Python, including the ability to read, debug, and patch an existing production codebase
  • Experience building monitoring, alerting, logging, or tracing capabilities, rather than only operating within an established environment
  • Hands-on responsibility for scheduled data pipelines and their day-to-day reliability
  • Experience troubleshooting dependencies, retries, missing or late data, failed jobs, and pipeline reruns
  • Experience operating applications and data workflows in a cloud environment such as Azure, AWS, or GCPExperience running an LLM-backed application in production and troubleshooting the operational issues unique to LLM-based systems
  • Meaningful on-call or incident-response experience with the ability to independently own an issue from detection through resolution and postmortem
  • Comfortable operating independently within a small technology team
Nice to Have:
  • Experience at a fintech or AI product company, proprietary trading firm, broker-dealer, hedge fund, or other financial-services organization
  • Familiarity with Microsoft Graph API, Entra app registrations, Teams bots, SharePoint permissions, or the broader Microsoft 365 ecosystem
  • Understanding of trading systems, market data, OMS/EMS workflows, end-of-day processes, settlement, or Bloomberg/vendor dataFIX familiarityC++ exposure
Who This Role Is ForThis is a strong opportunity for an engineer who enjoys owning production systems rather than working on one narrow piece of a large infrastructure organization. The ideal candidate has been the person on call, has diagnosed real production incidents, has kept scheduled pipelines running, and is comfortable figuring out whether an issue sits in the application, infrastructure, data, or upstream dependency.

Also on the board Same function, level within a rung

Level

Lead

Location

New York, NY

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

3 days ago

Apply for this role →