Hotline:+2203533578

NOC Technician 1300 views

Job Overview

Role Purpose

The NOC Technician is responsible for the continuous monitoring, triage, and first-line troubleshooting of the organization’s network infrastructure, server systems, and critical services from the Network Operations Centre (NOC).

This role serves as the frontline defense against service degradation and outages, ensuring rapid detection, escalation, and resolution of incidents to maintain optimal uptime and performance across all MIDSA environments.

The NOC Technician also plays a key role in customizing and maintaining the Checkmk monitoring
platform, including dashboard development, alert tuning, and automation of operational workflows through DevOps practices, directly supporting the IT Infrastructure team’s mission to deliver reliable, secure, and high performing technology services.

Principal Duties and Responsibilities

  • Monitor all network devices, servers, applications, and services in real time via the Checkmk monitoring platform and supplementary tools, ensuring 24/7 visibility into infrastructure health and performance.
  • Perform first-line troubleshooting and diagnosis of network connectivity issues, server faults, service degradations, and hardware failures, applying systematic root-cause analysis to restore services within defined SLA targets.
  • Customize and maintain Checkmk dashboards, views, and visualizations to provide actionable, role-specific monitoring insights for the NOC, infrastructure engineers, and management stakeholders.
  • Develop and refine Checkmk alerting rules, notification policies, and escalation paths to minimize false positives, reduce alert fatigue, and ensure critical incidents are flagged and routed to the appropriate resolver groups.
  • Author and maintain automation scripts (Bash, Python, or PowerShell) and DevOps pipelines to automate Checkmk agent deployments, configuration management, host discovery, and monitoring-as-code practices.
  • Manage and troubleshoot LAN, WAN, VLAN, VPN, firewall, and wireless infrastructure components, including switches, routers, access points, and security appliances across all MIDSA sites.
  • Escalate complex or high-severity incidents to Level 2/3 engineering teams with comprehensive diagnostic data, timeline documentation, and preliminary root-cause hypotheses to accelerate resolution.
  • Maintain and update the NOC runbook, standard operating procedures (SOPs), and knowledge base articles for recurring incidents, ensuring documentation is current, accurate, and actionable.
  • Perform routine systems administration tasks including patch verification, log analysis, backup monitoring, certificate expiry tracking, and capacity utilization reviews across Windows and Linux environments.
  • Collaborate with DevOps and infrastructure engineering teams to integrate monitoring into CI/CD
    workflows, implement infrastructure-as-code (IaC) monitoring definitions, and validate deployment health checks.
  • Conduct periodic health checks, performance baseline assessments, and trend analysis on critical
    infrastructure components, generating reports that inform capacity planning and proactive maintenance decisions.
  • Participate in incident post-mortems and problem management reviews, contributing technical findings and recommending preventive measures to reduce recurrence of known failure modes.
  • Support the onboarding of new infrastructure assets into the monitoring ecosystem, including agent installation, service discovery configuration, threshold calibration, and validation testing.

Technical Competency

  • Strong hands-on expertise in network troubleshooting across TCP/IP, DNS, DHCP, SNMP, HTTP/S, and routing/switching protocols, with the ability to diagnose issues using tools such as Wireshark, traceroute, ping, nmap, and netstat.
  • Proficiency in Checkmk administration, including host and service configuration, custom check
    development, Business Intelligence (BI) aggregations, agent bakery management, and WATO/Setup interface navigation.
  • Demonstrated DevOps capability, including scripting in Bash, Python, or PowerShell for automation, and familiarity with version control (Git), CI/CD pipelines, and infrastructure-as-code tools (Ansible, Terraform, or similar).
  • Solid understanding of Windows Server and Linux (Ubuntu, CentOS/RHEL) administration, including service management, log analysis (journalctl, Event Viewer), filesystem management, and user/permission administration.
  • Experience with firewall and security appliance management (pfSense, FortiGate, or similar), including rule configuration, NAT, VPN setup, and traffic analysis.
  • Knowledge of virtualization platforms (VMware vSphere, Proxmox, or Hyper-V) and container technologies (Docker) for monitoring integration and lab environments.
  • Familiarity with ITIL service management principles, particularly Incident Management, Problem
    Management, and Change Management workflows.
  • Experience with ticketing and ITSM platforms (e.g., ServiceNow, Jira Service Management, or GLPI) for incident logging, tracking, and SLA compliance reporting.
  • Strong analytical and reporting skills, with the ability to interpret monitoring data, identify trends, and produce clear operational and management reports.
  • Effective verbal and written communication skills for incident documentation, shift handover reporting, and cross-team technical collaboration.

Personal Competency

  • Strong attention to detail and situational awareness, with the ability to detect subtle anomalies in
    monitoring dashboards and system behaviour.
  • Calm and composed under pressure, capable of maintaining focus and systematic troubleshooting
    discipline during high-severity incidents and service outages.
  • Excellent problem-solving and analytical thinking skills, with a methodical approach to root-cause analysis and fault isolation.
  • Proactive and self-motivated, with a continuous improvement mindset and willingness to learn new technologies, tools, and operational practices.
  • Strong interpersonal and teamwork skills, able to collaborate effectively across shifts, teams, and
    organizational boundaries.
  • Reliable and punctual, with flexibility to work rotating shifts, weekends, and on-call schedules as required by NOC operations.
  • Strong organizational skills and the ability to manage multiple concurrent incidents and tasks while maintaining accuracy and prioritization.

Education

  • Diploma or Bachelor’s degree in Computer Science, Information Technology, Network Engineering, or a related technical discipline.
  • Minimum of 2-3 years of hands-on experience in a NOC, IT operations, or infrastructure support role within an enterprise environment.
  • Practical experience with network monitoring platforms, with specific experience in Checkmk being highly advantageous.
  • CompTIA Network+ or equivalent networking certification required; CCNA (Cisco Certified Network Associate) is highly advantageous.
  • Linux certification (LPIC-1, CompTIA Linux+, or RHCSA) is advantageous.
  • Checkmk Certified Engineer or equivalent monitoring platform certification is a strong advantage.
  • Demonstrable scripting and automation experience (Bash, Python, or PowerShell) in an operational or DevOps context.
Application should be submitted to apply@gamjobs.com and copy to jobs@gamjobs.com 
Note: Kindly mention the job title “NOC Technician” in the subject line of your application email.
Closing date 10th August, 2026.

Leave your thoughts

Company Information

Contact Us

https://gamjobs.com/wp-content/themes/noo-jobmonster/framework/functions/noo-captcha.php?code=777c7

Contact us

OIC Road, Old Yundum, Opp Swami India Building. Tel 3533578, info@gamjobs.com