Back to all jobs
M

Senior Solutions Architect – AI Infrastructure Security (East Coast)

Mirantis

Remote · Boston · Massachusetts · us Full-time 1h ago

Job description

About the role We are seeking a Senior Solutions Architect with deep security expertise to join the Voyager team in the Mirantis Office of the CTO. You will own the security solutioning of k0rdent AI, our platform for building and operating GPU clouds and AI factories, from metal to model. You will define how k0rdent AI environments are secured end to end: the hardware and firmware they boot from, the hosts, virtual machines and Kubernetes clusters they run, the networks and storage they share, the secrets and certificates that hold them together, and the AI models and applications they serve. You will turn that design into published reference architectures, working code and proofs of concept that run on real hardware with our customers and partners. As part of the Office of the CTO, research is a core part of the job. You will explore threats, designs and technologies before customers ask for them, prototype ideas that may not ship, and challenge established practice when a better approach exists. Your work will shape where k0rdent AI goes next, not only how it is deployed today. This role combines research with hands-on, outward-facing work. You will split your time between exploring new designs, writing architecture, building and validating it in the lab, and explaining it to engineers, security teams, executives and conference audiences. You will work with Solutions Architects, Partner Management and Engineering colleagues across many countries and time zones. Key Responsibilities Reference architecture Design and publish security reference architectures and solution designs for k0rdent AI, covering single-tenant private AI clouds, multi-tenant GPU clouds and hybrid deployments that span public cloud and on-premises infrastructure. Define a defence-in-depth model across the stack, from hardware root of trust and firmware to hosts, virtual machines, Kubernetes, networks, storage, models and applications. Define tenant isolation boundaries and their guarantees, and document where each boundary is strong, where it is weak, and what it costs to strengthen. Map designs to the controls and frameworks customers must meet, and document trade-offs (risk, cost, performance, operability) with defensible reasoning a customer can follow. Platform and infrastructure security Bare-metal and host security: secure and measured boot, TPM-based attestation, BMC and Redfish hardening, firmware integrity, and hardened Linux hosts. Virtual machine security: hypervisor and KubeVirt hardening, VM isolation, and the security implications of PCI passthrough and SR-IOV for GPUs and NICs. Kubernetes security: RBAC, Pod Security Standards, admission control and policy-as-code, runtime security, and secure multi-cluster and multi-tenant operation. Secrets management: secret stores, KMS and HSM integration, envelope encryption, secret delivery to workloads, and rotation. Certificates and PKI: internal certificate authorities, automated issuance and rotation, mTLS between services, and workload identity. Network security: segmentation and zero-trust designs, Kubernetes network policy, tenant isolation on front-end and GPU fabrics, and DPU-based security enforcement. Storage security: encryption at rest and in transit, key management, and tenant isolation on shared block, file and object storage used for training data and model weights. Public cloud security: secure landing zones, identity and access management, and network controls on major public clouds, and how they extend to hybrid deployments. AI model and application security Define how Transformer-based models are protected through their lifecycle: provenance and signing of model weights, integrity of training and fine-tuning data, and secure model registries. Secure inference services and AI applications against threats such as prompt injection, data and model exfiltration, model theft and abuse of agentic tool access. Design isolation for shared GPU infrastructure (MIG, vGPU, passthrough) and assess confidential computing options for protecting models and data in use. Secure the software supply chain for platforms, containers and models: image signing, SBOMs, vulnerability management and provenance. Research and exploration Track and evaluate emerging security technologies, standards and threats, such as confidential computing on GPUs, remote attestation, post-quantum cryptography, model signing and AI agent security. Prototype alternative designs in the lab, including ones customers have not yet asked for, and measure them against current practice. Challenge established or customer-preferred designs when evidence points to a better option, and make the case with data. Publish findings as internal research notes and design proposals and, where appropriate, as external papers, blog posts or talks. Code and proofs of concept Build and run proofs of concept with customers and partners, on Mirantis lab hardware and on customer sites, and report results against agreed success criteria. Write automation, policies and tooling (Python, Go, Bash, Ansible, Helm, Kubernetes manifests, Terraform, policy-as-code) that make the reference architectures reproducible. Assess designs and deployments through threat modelling, configuration review and hands-on testing, and turn findings into design guidance. Customers, partners and community Act as the security subject matter expert in customer discovery, design reviews and architecture workshops, bringing thought leadership and new ideas rather than only reflecting current practice. Engage with customer security, risk and compliance teams, and support security questionnaires and architecture reviews for strategic opportunities. Work with hardware, software and security partners (e.g. NVIDIA, server OEMs, security vendors) on joint designs and validations. Innovate and feed research results back to Product, Engineering and the Mirantis security team, and help shape the k0rdent AI roadmap. Present at industry events, webinars and partner summits, and run hands-on workshops for technical audiences. Write technical content: reference architecture documents, solution briefs, blog posts and internal enablement. Location, travel and working model This role is remote and based in the United States, with a strong preference for the East Coast. You will work with a team spread across many time zones; some meetings will fall outside standard local hours. Travel of up to 25% for customer engagements, partner meetings, lab work and industry events, primarily within the US and EU. Mirantis is an equal opportunity employer. Compensation and benefits are set according to the local market and employment arrangement. Qualifications Education and experience Bachelor's degree in Computer Science, Computer Engineering, Cybersecurity or a related field, or equivalent practical experience. 8+ years in security engineering or security architecture, with at least 3 years securing cloud, Kubernetes or large-scale infrastructure platforms. Customer-facing experience as a solutions architect, pre-sales engineer, consultant or technical lead. Required technical skills Cloud security: designing and operating secure deployments on at least one major public cloud (AWS, Azure or GCP) and on private cloud or bare-metal infrastructure. Host and virtualisation security: Linux hardening (e.g. SELinux or AppArmor, CIS benchmarks), secure boot, TPM, and hypervisor and VM isolation, including KubeVirt. Kubernetes security: RBAC, Pod Security Standards, admission controllers and policy engines (e.g. Kyverno, OPA Gatekeeper), runtime security (e.g. Falco), and multi-tenancy patterns. Secrets and key management: tools such as HashiCorp Vault or OpenBao, External Secrets Operator, cloud KMS and HSMs. PKI and identity: certificate lifecycle with tools such as cert-manager, mTLS, OIDC and SSO, and workload identity (e.g. SPIFFE/SPIRE). Network security: segmentation, firewalling, zero-trust architectures and Kubernetes network policy (e.g. Cilium, Calico). Storage and data security: encryption at rest and in transit, key hierarchy design, and access control on shared storage. AI security: a working understanding of how Transformer-based models are trained, packaged and served, and the threats specific to them (e.g. OWASP Top 10 for LLM Applications, MITRE ATLAS). Software supply chain security: image signing (e.g. Sigstore/cosign), SBOMs, vulnerability scanning and SLSA concepts. Programming or scripting in at least one language (Python or Go preferred), and comfortable using Git, CI and infrastructure-as-code. How you work We hire for these behaviours as much as for technical depth, and will ask for concrete examples of each. Curious. You dig into how attacks actually work, read the specs, advisories and source, and keep up with how AI and infrastructure threats are changing. You question established designs, including the ones customers are used to, when there is a better way. Ownership. You take a reference architecture or a POC from the first sketch to a published, validated result, and you stand behind it. Proactive. You spot risks and gaps in the product, the documentation or a customer design before you are asked, and you act on them. Fast learner. You get productive quickly in unfamiliar hardware, software or customer environments. Business acumen. You frame security decisions in terms of risk, cost, compliance and time-to-value, and explain them in those terms to decision makers. Communication and collaboration Excellent written and spoken English; you write documents others can build from. Comfortable presenting to large audiences at conferences and events, and running hands-on workshops. Able to adapt the message to the audience, from engineers to CISOs and customer executives. Experience working in an international, distributed company, across time zones and different work cultures, with Solutions Architects, Partner Management and Engineering teams. Nice to have Experience with compliance frameworks and benchmarks relevant to cloud and AI infrastructure, such as NIST SP 800-53, NIST AI RMF, FedRAMP, SOC 2, ISO/IEC 27001, CIS benchmarks or DISA STIGs. Hands-on experience with confidential computing (e.g. AMD SEV-SNP, Intel TDX, NVIDIA GPU confidential computing) and remote attestation. Experience securing GPU cloud, neocloud or HPC environments, or NVIDIA Cloud Partner (NCP) style deployments. Experience with DPUs (e.g. NVIDIA BlueField) for infrastructure isolation and security offload. Offensive security or red-teaming experience, including against AI systems. Experience with Mirantis products, MKE, k0s, k0rdent security. Security certifications such as CISSP, CCSP, CKS, OSCP, or public cloud security specialty certifications. Contributions to open-source security or Kubernetes projects, or a track record of public talks and technical writing. Participation in standards or industry bodies (e.g. CNCF TAG Security, OpenSSF, Confidential Computing Consortium, CoSAI), or published research, CVEs, patents or white papers. Additional languages beyond English. Why you’ll love Mirantis Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises Collaborate with a world-class, distributed team committed to openness and technical excellence Shape the product narrative and influence go-to-market success It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities. #remote We are a Leader for Container Management in G2 (#2 after AWS)! About Mirantis Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Similar open jobs