We are seeking a technically strong Application Engineer who brings a software engineering mindset to production support. This role focuses on ensuring application stability and reliability while driving improvements through automation, observability, and modern Site Reliability Engineering (SRE) practices. The ideal candidate will be responsible for monitoring production systems, managing incidents, and performing code-level troubleshooting. In addition, they will contribute to building automation solutions, improving monitoring frameworks, and leveraging AI-assisted tools to enhance operational efficiency and reduce manual effort.
Key Responsibilities:
Incident Management
Act as the primary point of contact for production issues, alerts, and user-reported incidents
Log, categorize, and prioritize tickets in accordance with defined SLAs
Ensure timely resolution and minimize business impact
Triage and Troubleshooting
Perform root cause analysis for production issues across application and infrastructure layers
Collaborate with cross-functional teams to resolve complex technical problems
Use logs, metrics, and traces to diagnose issues efficiently
Code-Level Analysis and Fixes
Review and debug application code to identify and resolve defects
Implement code-level fixes and coordinate with development teams for larger changes
Participate in code reviews to ensure quality and maintainability
Automation and Operational Efficiency
Design and build automation scripts and tools to reduce manual operational tasks
Leverage AI-assisted tools to streamline troubleshooting and incident response
Continuously identify opportunities to improve operational workflows
Observability and Monitoring
Build and maintain dashboards, alerts, and monitoring frameworks
Enhance system observability to proactively detect and prevent issues
Ensure monitoring coverage aligns with business-critical services
SRE and Continuous Improvement
Apply SRE principles to improve system reliability, availability, and performance
Participate in post-incident reviews and drive corrective actions
Contribute to capacity planning and performance tuning efforts
Escalation and Collaboration
Escalate unresolved issues to appropriate teams while maintaining ownership of resolution
Work closely with development, QA, and infrastructure teams to prevent recurring issues
Foster strong cross-team relationships to support rapid issue resolution
Documentation and Knowledge Management
Create and maintain runbooks, troubleshooting guides, and knowledge base articles
Document incident details, root causes, and resolutions for future reference
Ensure documentation remains current and accessible to the team
Communication
Provide clear and timely updates to stakeholders during incidents
Communicate technical issues effectively to both technical and non-technical audiences
Collaborate with global teams across different time zones
Required Qualifications:
Experience
3+ years of experience in application support, production support, or a similar engineering role
Technical Skills
Strong understanding of software engineering principles and ability to read and debug code
Experience with monitoring and observability tools
Familiarity with cloud platforms and containerized environments
Automation and DevOps
Experience building automation scripts using Python or similar languages
Familiarity with CI/CD pipelines and DevOps practices
Soft Skills
Strong problem-solving and analytical skills
Excellent verbal and written communication skills
Ability to work effectively under pressure during incidents
Education
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience
Preferred Qualifications:
Experience with AI-assisted operational tools
Exposure to Site Reliability Engineering (SRE) practices
Experience working in a 24/7 production support environment