Required Qualifications:
• 7+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
Desired Qualifications:
• Mainframe technologies support/operations experience or exposure.
• Technical experience/exposure working with Mainframe technologies, etc. and basic IBM utilities.
• Understanding of payment processing on a large scale.
• Experience/Exposure using monitoring tools (example: Splunk, AppDynamics, Grafana, Geneos ITRS etc.)
• Familiarity with ServiceNow for incident, problem, and change management (opening incidents, reviewing changes, documenting records).
• Excellent leadership, verbal, written, and interpersonal communication skills.
• Experience with change management.
• Understanding of Site Reliability Engineering (SRE) principles, with hands-on experience applying them
• Experience with Mainframes with below technical skills
• Operating Systems
• MVS z/OS
• Windows OS
• Programming / Languages
• COBOL
• JCL
• CICS
• Databases
• DB2
• VSAM
• IMS DB
• Stored Procedures
• Tools
• SPUFI
• QMF
• PLATINUM
• XPEDITOR
• Version Control
• ENDEVOR
• Scheduler
• CA‑7
• Application Services
• MQ
• NDM
Job Expectations:
• Forward thinking and innovative to identify and implement best practices/procedures.
• Ability to independently run incident recovery calls (incident management) and review changes to mitigate risk.
• Coordination with Vertical Application Support Teams:
• Partner with application owners and support teams to understand system architecture, performance baselines, and critical business transactions.
• Process improvement:
• Optimize alerts, incident reductions, and reports to provide actionable insights.
• Automation and Scripting: This includes scripting using REXX.
• Develop automation scripts using REXX to avoid manual work.
• SQL and Data Analysis
• Use SQL to extract, analyze, and correlate data from monitoring platforms and application databases.
• Support root cause analysis and performance tuning through data-driven insights.
• Incident Response and Continuous Improvement: This includes managing incidents and changes effectively.
• Participate in incident triage and post-mortem analysis to identify observability gaps. This includes managing incidents and changes effectively
• Continuously refine monitoring strategies based on feedback and evolving application landscapes.
• Documentation and Knowledge Sharing:
• Maintain clear documentation of observability configurations, standards, and best practices.
• Conduct knowledge-sharing sessions with support teams to promote observability and maturity.
Posting End Date:
24 Sep 2026*Job posting may come down early due to volume of applicants.