About the position
ENVIRONMENT:
Our client operates a real-time communications platform offering call tracking, number privacy, click-to-call, SMS and WhatsApp solutions to clients across South Africa, the UK and Europe. When a call has to connect, it has to connect. They are hiring one person to keep the platform healthy, including the infrastructure it runs on, its security posture and the customers who depend on it. This is a hands-on role offering real ownership from day one, working directly with a small, senior engineering team who will provide training on the technology stack.
You will own three key areas:
Platform & DevOps (~40%)
- Deploy and monitor their production Docker Swarm fleet on AWS, including EC2, RDS PostgreSQL, ElastiCache Redis, S3 and SES.
- Keep the Celery queue environment healthy by identifying backlogs, stalled schedulers and memory-intensive workers before customers are affected.
- Manage their ELK stack, including Elasticsearch, Logstash, Kibana and Metricbeat, as well as their alerting rules.
- Triage Sentry errors, own the weekly systems-health review and keep runbooks up to date.
- Support continuous integration through GitHub Actions, staging deployments and database maintenance.
Security (~25%)
- Own the SentinelOne EDR console, including endpoint coverage, alert triage and escalation.
- Run Greenbone vulnerability scans, act on the findings and ensure scan targets remain accurate.
- Own GitHub security hygiene, including Dependabot alerts, dependency upgrades, secret scanning and branch protection.
- Maintain fail2ban, AWS WAF and firewall rules across the voice and Session Border Controller (SBC) fleet.
- Track penetration-testing findings through to closure.
Product & Customer Support (~35%)
- Provide Tier 1–2 support for their Voice, SMS and WhatsApp products.
- Investigate call-quality, routing and Call Detail Record (CDR) issues across FreeSWITCH, FusionPBX and their SBCs.
- Investigate SMS delivery and Delivery Receipt (DLR) issues with SMPP suppliers.
- Provision numbers (DIDs), configure routing and transfer numbers between accounts.
- Log and drive supplier tickets while keeping customers and account managers informed.
What They Need from You:
- At least two years of experience in a commercial DevOps, Systems Administration, Network Operations Centre (NOC), Site Reliability Engineering (SRE) or technical-support engineering role. This is not an entry-level position.
- Confidence working with Linux (Ubuntu), including shell, Systemd, journald, networking and SSH.
- Practical Docker experience, with the ability to read a Compose file and troubleshoot a container that will not start.
- Working knowledge of AWS, including EC2, RDS, S3, IAM and security groups.
- Sufficient Python knowledge to read Django code, run a management command and fix a minor bug.
- SQL skills sufficient to answer support-related questions using the database.
- Confidence working with Git and GitHub.
- Willingness to carry an on-call phone as part of a shared rota.
- Strong written communication skills. A significant part of the role involves clearly explaining what happened and what was done to resolve it.
On-Call Responsibilities:
The platform carries live calls and messages and therefore requires support outside standard office hours. You will join a shared after-hours rota with the engineering team. You will never be the only person available, as every escalation path ends with a senior engineer. You will not join the rota until you can perform production deployments without supervision.
A standby allowance applies for every week that you carry the phone, and any night work is compensated with leave.
Bonus Points For:
- Exposure to VoIP/SIP technologies such as FreeSWITCH, Asterisk or Kamailio, or experience with SMPP/SMS.
- Experience with Elasticsearch, Kibana, Grafana or similar observability tools.
- Experience with security tools and practices, including EDR, vulnerability scanners, WAF and CIS hardening.
- Experience with Ansible, Terraform or other Infrastructure-as-Code tools.
- A security certification such as Security+, eJPT or an Azure/AWS security certification, or a Linux certification such as RHCSA or LFCS.
Desired Skills:
- Platform
- Support
- Engineer