windows serverlinuxvmwarecloudawsazuregcppowershellbashpythonansibleterraformdevopsci/cddockerkubernetessite reliability engineeringmonitoringservice nowjiraautomationinfrastructuresreincident managementitilactive directorynetworkingsecurityvirtualization
Key details
Salary
Not specified
Employment type
Full-time
Seniority
Mid-level
Years experience
3-5
Location
Bengaluru, India
Full job description
The Infrastructure & Reliability Engineer will operate, support, secure, and improve enterprise infrastructure across Windows, Linux, virtualized, cloud, and hybrid environments. Responsibilities include L2 systems administration, L1 site reliability engineering practices, incident and change management, monitoring, automation, and collaboration with various teams. Required skills include 3-5 years experience in Windows/Linux administration, cloud operations, scripting/automation, incident management, and monitoring. Preferred qualifications include a relevant degree, exposure to cloud platforms (AWS, Azure, GCP), virtualization, container technologies, SRE concepts, observability tools, databases, and certifications. The role requires participation in shifts and on-call support. The position is located in Bengaluru, Karnataka, India.
What you'll do
Administer, troubleshoot, patch, upgrade, and support Windows Server and enterprise Linux platforms
Support Windows services including Active Directory, DNS, Group Policy, WSUS, permissions, access, and service accounts
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Perform Linux administration activities including user management, permissions, file systems, services, packages, processes, logs, networking, storage, and performance troubleshooting
Support VMware, Proxmox/KVM, physical servers, and cloud infrastructure (AWS, Azure, GCP)
Execute routine health checks, capacity reviews, backup validation, vulnerability remediation, configuration fixes, and service restoration
Diagnose infrastructure issues across OS, storage, compute, network, identity, cloud, virtualization, and security layers
Support enterprise server hardware using vendor management interfaces and diagnostic tools
Manage incidents, service requests, and change records using ITIL-aligned platforms like ServiceNow or Jira
Restore services within SLA targets through troubleshooting, escalation, communication, and documentation
Participate in high-priority incident response, evidence collection, RCA, and post-incident reviews
Execute approved changes with validation, rollback planning, and stakeholder communication
Identify recurring issues and contribute to problem management and operational improvements
Coordinate with application, cloud, security, database, network, and vendor teams for issue resolution
Monitor availability, performance, service health, capacity, and reliability indicators using enterprise monitoring platforms
Triage alerts, validate impact, reduce alert noise, escalate risks to senior teams
Support service reliability reviews and apply foundational SRE practices
Assist with resilience, failover, recovery, and operational-readiness testing
Recommend automation and improvements to reduce repetitive manual activities
Automate routine tasks using PowerShell, Bash, Python, Ansible, Terraform, or workflow tools
Use Git-based version control and peer review for scripts and infrastructure-as-code
Support CI/CD and DevOps operational activities using Jenkins, GitHub, Bitbucket, Argo CD, Buildkite
Integrate REST APIs and workflow tools to improve operational efficiency
Support basic Docker and Kubernetes troubleshooting and escalate as needed
Implement security hardening, remediate vulnerabilities, support audits, and maintain operational evidence
Maintain SOPs, runbooks, knowledge articles, configuration records, and documentation
Follow security policies, access controls, change controls, and compliance processes
Support patch compliance, endpoint protection, backup validation, vulnerability closure, and risk reduction
Requirements
3 to 5 years hands-on experience in Windows/Linux systems administration, infrastructure operations, production support, cloud operations, DevOps, or SRE
Strong working knowledge of Windows Server administration and/or Linux administration
Experience supporting enterprise infrastructure including virtualization, networking, storage, backup, identity, security, monitoring, and cloud fundamentals
Practical scripting or automation experience using PowerShell, Bash, Python, Ansible, Terraform, or equivalent tools
Experience with incident, request, change, problem, and knowledge-management processes in an enterprise environment
Experience with monitoring, logging, dashboards, alert triage, performance analysis, and capacity management
Ability to troubleshoot complex infrastructure issues and communicate clearly during service-impacting events
Ability to participate in rotational shifts, weekend support, maintenance windows, and on-call support
Preferred: Bachelor’s degree in Computer Science, IT, Engineering or equivalent experience
Preferred: Exposure to AWS, Azure, GCP, VMware, Proxmox/KVM, Docker, Kubernetes, and hybrid infrastructure
Preferred: Foundational understanding of SRE concepts including SLIs, SLOs, error budgets, toil reduction, capacity planning, and blameless post-incident review
Preferred: Familiarity with observability platforms such as Splunk, ELK/Kibana, Grafana, Prometheus, CloudWatch, Datadog, New Relic
Preferred: Familiarity with databases such as MySQL or PostgreSQL
Preferred: Relevant certifications in Microsoft, Linux/Red Hat, VMware, AWS, Azure, GCP, ITIL, DevOps, or SRE practices
Employee well-being focusCollaborative work environmentOpportunities for growth, learning, development, and career advancementInnovation-driven cultureWork-life balance and flexibilityDiversity, inclusion, and equal employment opportunity commitment
Apply now
Ready to take the next step in your career? Click the button below to continue to the application process.