S
Senior Agent Operations & Evaluation Engineer #AIDA
Singapore Telecommunications Ltd
Singapore · sg
4h ago
99%
Strong
Job description
Powering the Future with AIDA To lead the next phase of our AI evolution, we’ve launched a new business unit AIDA – Artificial Intelligence & Data Analytics – a strategic engine driving our transformation designed to scale our AI ambitions with precision and purpose. This marks a pivotal shift in how we operate, innovate, and serve to embed intelligence into every layer of our business. At Singtel, this is more than a technology upgrade. It’s a strategic transformation that redefines how value is created across the enterprise core—augmenting human capabilities and unlocking entirely new potential. It is a transformation journey by aligning people, platforms, and processes under one cohesive strategy. Our mission is to build AI literacy and foster a culture where intelligence empowers people. We welcome you to join us on a transformational journey that’s reshaping the telecommunications industry — and redefining what’s possible with AI at its core. Grow with us in a workplace that champions innovation, embraces agility, and puts human potential at the heart of everything we do. About the Role: An engineering role that operates the live GenAI agents and builds the evaluation and observability that proves they are accurate, safe and improving. On the operations side, keeps agents healthy, performs first-level triage and incident response, and minimises downtime. On the evaluation and observability side, builds the operational dashboard for each use case, defines the metrics that matter, runs the evaluation suites, and proactively detects issues before customers do. Works closely with the Day 1 build teams, solution architects and business owners across Singtel Singapore. How You will Make An Impact: Build and maintain the operational observability dashboard for each AI / agent and Classical AI ML use case, across system health, performance, quality, safety and business-outcome metrics. Define, with build teams and solution architects (technical metrics) and product / business owners (outcome metrics), what needs to be measured per use case; consume the telemetry the platform provides. Build and run the evaluation suites (offline before launch and continuous evaluation on sampled production traffic after) and drive proactive detection of drift and quality decay. Monitor live agents, using the dashboards and automated safeguards, and act on alerts and proactive signals. Perform first-level triage and in-scope fixes (restart, safe-listed configuration, cache reset, routing fallback) per the runbooks. Manage incidents — classify severity, route by fault domain, escalate with a full evidence bundle, and verify-and-close in production — while owning the incident throughout. Contribute to the handover gate — verifying operability and evaluation evidence before an agent is accepted into full operation. Skills for Success: Degree in Computer Science, Engineering, Data Science, or a related field 6 - 10 years across software / SRE / data or ML engineering Hands-on production operations experience AND hands-on evaluation, data-analysis or quality experience Experience with LLM / agentic systems Monitoring, alerting and observability; log / trace analysis Evaluation methods and metric design for AI systems Dashboarding tools and Python Understanding of LLM / agent architectures and failure modes LLM observability and evaluation tooling (e.g. tracing, eval harnesses) Incident management practice Composure and clear communication under incident pressure Analytical rigour and attention to detail Ability to context-switch and prioritise across two disciplines Clear technical writing for gate findings Ability to influence build teams on quality Understanding of telco customer intents and journeys Are you ready to say hello to BIG Possibilities? Join Singtel to shape what's next and accelerate your career through meaningful work, continuous learning, and real impact.