Mid-Level AWS Production Support EngineerDirect Full-TimeLafayette, LA | Knoxville, TN | Columbia, SC | Birmingham, ALPosition Summary We are seeking an AWS Production Support Engineer with 2–5 years of production/application support experience to provide 24x7 operational support for applications running in an AWS environment. The ideal candidate will have strong troubleshooting and incident-management skills, hands-on experience investigating application and infrastructure alerts, and the ability to perform initial triage using AWS services, Splunk, OpenTelemetry, and application logs. This role requires someone who is proactive, self-motivated, comfortable working in a production environment, and able to quickly determine whether an issue is isolated or part of a broader outage.Key Responsibilities
- Provide 24x7 production support for applications operating in an AWS environment.
- Monitor production alerts and investigate application or infrastructure incidents.
- Perform initial troubleshooting and triage using AWS, Splunk, OpenTelemetry, and application/system logs.
- Investigate batch and scheduled job failures, gather relevant information, and identify potential causes.
- Determine the severity and impact of production issues and escalate to appropriate SMEs or engineering teams when required.
- Troubleshoot application performance, availability, and operational issues.
- Support production deployments and releases.
- Execute and monitor basic CI/CD jobs and deployment processes.
- Review logs and monitoring data to determine whether an issue is isolated or part of a larger production outage.
- Work closely with development, infrastructure, cloud, and support teams during incident resolution.
- Maintain and continuously improve operational documentation and knowledge articles.
- Identify opportunities to improve production-support processes and operational efficiency.
Required Qualifications- 2–5 years of Production Support / Application Support experience.
- Experience supporting applications in a 24x7 production environment.
- Working knowledge of AWS and the ability to navigate AWS services for troubleshooting and investigation.
- Hands-on experience with Splunk for log analysis and troubleshooting.
- Strong experience with incident investigation, initial triage, troubleshooting, and escalation.
- Experience investigating application alerts, system logs, and production failures.
- Experience supporting and troubleshooting batch jobs / scheduled jobs.
- Strong analytical and problem-solving skills.
- Ability to determine whether an issue is isolated or part of a broader outage.
- Ability to work independently and escalate appropriately when deeper technical expertise is required.
- Strong communication skills and a proactive, self-learning mindset.
Preferred / Nice-to-Have Skills- OpenTelemetry
- AutoSys
- Apache
- RMJ
- CI/CD
- Production release support
- Application debugging
- Monitoring and observability tools
- Knowledge-base / operational documentation experience
Work Schedule This position supports a 24x7 production environment and requires flexibility to work rotating 12-hour shifts. Typical schedule: 3.5 days on / 3.5 days off, including a Sunday–Wednesday rotation, depending on production-support coverage requirements. Candidates must be comfortable working flexible schedules as required to support production operations. Ideal Candidate Profile The successful candidate will be a hands-on production-support professional who: - Enjoys troubleshooting production issues.
- Can investigate alerts independently before escalating.
- Is comfortable working with ambiguity during operational incidents.
- Learns new applications and technologies quickly.
- Understands incident severity and escalation procedures.
- Can work effectively under pressure during production outages.
- Continuously looks for ways to improve support processes and documentation.
Interview Focus Areas Candidates should be prepared to discuss: - A production incident they handled and their troubleshooting approach.
- How they investigate issues using AWS logs or Splunk.
- Steps taken when a batch or scheduled job fails.
- How they determine when an issue should be escalated.
- Experience supporting a 24x7 production environment.
- Examples of operational-process or documentation improvements they have implemented.
#M1 #DI-CB2 #L1 - KB1
Ref: #404-IT Pittsburgh