F
Senior Platform Reliability Engineer (Fabric and Interconnect)
Firmus Technologies
Remote · Sydney · au
10h ago
78%
Strong
Job description
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. AI FactoryOS Operations AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next. Role Summary Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Platform Reliability Engineer, Fabric and Interconnect, owns the reliability of the fabrics this estate runs on: the GPU-to-GPU interconnect domains, the high-performance network fabrics carrying training and inference traffic, and the DPU-based host networking that binds compute to the rest of the platform. This is a hands-on senior role with deep technical expertise . Automation is a first-class part of the role: the team builds and maintains the guarded automation and remediation tooling that turn manual fabric response into a self-healing capability, and the role engages fabric vendors at engineering level, reproducing faults to their standard and holding them to their answers. Key Responsibilities Responsible for the reliable operation, automation and continuous improvement of the estate's GPU interconnect and network fabrics (for example NVLink and NVSwitch domains, InfiniBand and Spectrum-X). Build and maintain the guarded automation and remediation tooling for fabric faults, contributing to the software-driven remediation of AI clusters, including fault isolation and fabric reconvergence. Diagnose and tune performance across the interconnect stack, from application collective communication down to link level, working with technologies including NVLink , InfiniBand, RoCE and congestion control tuning. Operate DPU-based host networking across the fleet, including offload path configuration and driver and firmware compatibility. Execute firmware upgrade waves, fabric expansions and capacity changes to the supported paths and scaling patterns defined by AI Infrastructure, owning the production window, the staged or canary path, verification against declared success criteria, and rollback execution. Provide the deepest technical expertise for fabric and interconnect faults, correlating a collective communication failure to a specific physical link and diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure . Lead vendor escalations at engineering level, reproducing faults to the vendor's standard and pushing back credibly when a diagnosis does not explain the observed behaviour . Lead technical recovery during major fabric incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results. Skills & Experience Strong skills in high-performance networking and systems engineering, with 8+ years of experience including substantial ownership of production network or interconnect infrastructure in a 24/7 environment. Deep operational experience with high-performance GPU interconnect fabrics (for example NVLink and NVSwitch ), including domain topology and failure diagnosis. Extensive experience with high-performance networking fabrics (for example InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X), including routing internals and congestion control tuning. Experience with DPU or SmartNIC -based host networking, including offload paths and driver and firmware coordination. Experience planning and executing firmware upgrade waves and capacity expansions on production fabrics with defined rollback. Strong skills in infrastructure automation and infrastructure-as-code practices, with change delivered through peer review and progressive rollout. Practical experience with scripting or programming for operational automation and tooling, such as Python, Go or Bash. Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, vendor escalation at engineering level, and the production of runbooks that others can execute successfully. Clear technical judgement and communication skills, with the ability to explain complex fabric failures to engineers and non-specialists. Preferred Experience Experience operating fabrics for large-scale distributed training or inference workloads. Experience with NVIDIA rack-scale or multi-node GPU systems and their interconnect topology. Experience in a multi-tenant service provider, cloud or colocation environment. Knowledge of data centre and hardware fundamentals, including cabling and optics, firmware management and hardware fault workflows. Relevant vendor certification . Location & Reporting Location : Based in Australia or Singapore, with travel to Australian AI Factory sites as required. On-call : The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function. Reporting to : Reports to the Head of AI FactoryOS Operations while the function is being established , working under broad direction with a high degree of autonomy and direct access to the decision makers. As the function reaches its planned structure, the role will report to the Infrastructure Operations Manager, with the Head of AI FactoryOS Operations remaining accountable for the function. The scope, level and remit of the role do not change under either arrangement. Employment Basis Permanent full-time Diversity At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions. Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.