Back to all jobs
B

Principal Machine Learning Engineer

Blazetalent

NY · us Full-time 2d ago

Job description

Principal ML Engineer New York, NY (Hybrid) About the Company We're building AI-native enforcement infrastructure for enterprise communication — technology that catches and fixes compliance issues in real time, before an AI-generated message ever reaches a customer or counterparty, across every channel where AI represents the business. Most existing tools only flag problems after the fact, once the risk is already out the door; we intervene before send. This is a new category, and we're the ones defining it. We're backed by top-tier venture capital and built by a team with backgrounds at major tech and financial firms, led by a founder who has built and scaled AI companies before. The Role Specialized language models sit at the core of our enforcement layer, making real-time decisions about whether a communication is safe to send. These models need to be accurate, fast, and dependable, since they operate directly in the path of live traffic. Our research team owns the underlying science — model behavior, training objectives, data strategy, and quality standards. You'll own the systems that turn that science into a reliable, production-grade product: the pipelines that train models reproducibly, the evaluation infrastructure that proves they work, and the serving stack that runs them at scale. This is a hands-on, principal-level individual contributor role on a small, senior team. It's a systems and infrastructure role, not a research role — ideal for someone who loves making ML industrial-grade. What You'll Do Build and own training pipelines: data prep, reproducible fine-tuning runs, experiment tracking, and release automation Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management Ship model updates safely with versioning, canarying, rollback, and drift monitoring Build repeatable workflows for adapting models to new domains and customer needs Convert expert labels and reviewer feedback into clean training and evaluation datasets Set the technical bar for ML infrastructure as the team grows What We're Looking For 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production Hands-on depth with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM Experience building eval harnesses, regression gates, or dataset pipelines, with solid understanding of precision, recall, and calibration Proven track record owning model serving under real latency, reliability, and cost constraints — not just in notebooks Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability Comfort with high ownership on a small team: scoping your own work, shipping weekly, and making pragmatic build-vs-buy calls Enjoyment of close collaboration with a research counterpart, with clear interfaces and no turf wars Nice to Have Experience productionizing small or specialized language models Experience with structured-output serving or constrained decoding in production Background in a regulated or high-stakes domain such as fintech, healthcare, legal, or trust and safety Experience deploying models into customer-controlled environments Compensation & Benefits $200,000–$250,000 base salary, depending on experience Performance bonus and meaningful early-stage equity Health, dental, and vision coverage Hybrid work from a New York office

Similar open jobs