Job Description
Serve as a senior-level technical leader within IT Solutions Delivery, responsible for deploying, operating, and continuously improving production AI-enabled platforms and services that support critical business applications.
This role ensures that AI/ML capabilities are delivered into production environments using the same operational rigor, reliability standards, and support models as enterprise IT infrastructure, enabling consistent uptime, performance, and scalability. The engineer partners closely with application teams, platform engineering, and IT operations to ensure AI services are production-ready, supportable, and aligned to enterprise operational standards.
Job Duties and Responsibilities:
Deploy AI/ML solutions into enterprise production environments using repeatable, low-risk release processes
Build and maintain automated pipelines that support solution delivery across development, testing, and production
Ensure all AI services meet enterprise standards for deployment, configuration, and change management
Own day-to-day operations of AI-enabled platforms, ensuring availability, reliability, and performance of business-facing services
Establish and enforce site reliability engineering (SRE) practices, including high availability and fault tolerance, capacity planning and auto-scaling, and redundancy and failover strategies
Continuously optimize platform performance, resource utilization, and cost efficiency
Implement and maintain monitoring, logging, and alerting aligned to enterprise ITOM practices
Define and track service health, performance, and data quality metrics for AI-enabled services
Configure proactive alerting for degradations or anomalies and integrate with enterprise event management platforms
Develop and maintain operational dashboards and visibility tools for ongoing service assurance
Provide production support for AI-enabled services, including incident triage and resolution, performing root cause analysis (RCA), and implementing corrective and preventative actions
Develop and maintain runbooks and support procedures to enable consistent issue resolution
Partner with operations and service desk teams to ensure support readiness and knowledge transfer
Integrate AI services into enterprise application and infrastructure ecosystems, ensuring compatibility with existing platforms
Collaborate with solution delivery, infrastructure, and application teams to ensure services are fully operationalized, properly monitored, and supportable through standard IT processes
Ensure AI components behave as first-class enterprise services within the broader application landscape
Ensure all AI services are deployed and operated in alignment with enterprise security, compliance, and data protection standards
Manage the full operational lifecycle of AI-enabled services, including versioning and controlled releases, performance tuning and optimization, and continuous improvement of deployment and support processes
Identify opportunities to automate and standardize platform operations to improve efficiency and reliability
Serve as a subject matter expert in AI platform operations
Lead or support resolution of major incidents and complex operational challenges
Drive adoption of standardized operational practices, including runbooks, and reliability engineering
Provide guidance to project teams to ensure solutions are designed for production support from day one
Requirements
Education and Experience
Undergraduate degree in computer science, information technology, or related curriculum (or equivalent combination of experience and education)
8 or more years of information technology, cloud or platform engineering, DevOps, site reliability engineering (SRE), infrastructure engineering, application operations, or related experience that includes experience:
supporting AI/ML platforms, machine learning operations (MLOps), AI-enabled applications, large-scale data platforms, or other advanced analytics environments in production
deploying, monitoring, and supporting business-critical applications in cloud-based and hybrid enterprise environments
designing and managing CI/CD pipelines, automated deployment processes, and infrastructure-as-code solutions
partnering with application development, infrastructure, security, and operations teams to operationalize new technologies and services
serving as a technical lead, senior engineer, or escalation point for complex production issues
Certification and/or License - may be required during course of employment
Knowledge, Skills, and Abilities
Deep understanding of managing supported systems in a large-scale environment
Solid understanding of AI/ML operational practices, including model deployment, model monitoring, inference services, version management, and AI platform lifecycle management
Strong understanding of backup technologies and cloud technologies
Strong scripting and automation skills
Strong collaboration skills with application development, platform engineering, cybersecurity, infrastructure, and service desk teams
Strong problem solving and analytical skills with the ability to quickly isolate problems, collect data, establish facts, and draw valid conclusions; able to perform root cause analysis and implement sustainable corrective and preventive actions
Able to deploy and maintain highly available, scalable, and supportable AI-enabled services in production environments
Able to automate operational processes, platform provisioning, deployments, monitoring, and recovery activities
Able to serve as the senior technical escalation point for critical incidents and complex operational challenges
Able to influence teams and drive adoption of enterprise operational standards and best practices
Able to communicate complex technical concepts to both technical and non-technical stakeholders
Able to prioritize multiple operational demands in fast-paced production environments
Able to work independently with limited direction while maintaining accountability for enterprise-critical service
Must be able to read, write and speak English
An Equal Opportunity Employer including Disabled/Veterans
Pay Range $98000-$139000/Annually