Site Reliability Engineer - Service Assurance Systems
Viasat
Tanggal: 5 hari yang lalu
Kota: Batam, Riau
Jenis kontrak: Penuh waktu
About Us
One team. Global challenges. Infinite opportunities. At Viasat, we’re on a mission to deliver connections with the capacity to change the world. For more than 35 years, Viasat has helped shape how consumers, businesses, governments and militaries around the globe communicate. We’re looking for people who think big, act fearlessly, and create an inclusive environment that drives positive impact to join our team.
What You'll Do
The Service Assurance Systems (SAS) Group, part of Global Operations, develop and maintain many software systems and applications which support the operation of Viasat services.
The team collects assurance and operational data from a wide range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose, the team builds and maintains data pipelines that cleanse, transform and enrich raw data, making it readily consumable by our stakeholders across operations, engineering and management.
As a Site Reliability Engineer (SRE), you will play a key role in bridging the gap between software development and operational reliability. You will be responsible for supporting the applications built and maintained by the SAS group, managing deployments, investigating operational issues, and driving the team’s transition toward a modern DevOps culture. You will act as a first point of contact for service issues raised through the ServiceNow ticketing system, working to diagnose, triage and resolve incidents efficiently while collaborating closely with developers and operations teams.
The day-to-day
You will be working as part of a small team of developers and operations engineers, supporting the evolution of Viasat’s network and service monitoring capabilities, ensuring it remains world class in support of existing and future services.
Day-to-day the role will involve:
Viasat is proud to be an equal opportunity employer, seeking to create a welcoming and diverse environment. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, ancestry, physical or mental disability, medical condition, marital status, genetics, age, or veteran status or any other applicable legally protected status or characteristic. If you would like to request an accommodation on the basis of disability for completing this on-line application, please click here.
One team. Global challenges. Infinite opportunities. At Viasat, we’re on a mission to deliver connections with the capacity to change the world. For more than 35 years, Viasat has helped shape how consumers, businesses, governments and militaries around the globe communicate. We’re looking for people who think big, act fearlessly, and create an inclusive environment that drives positive impact to join our team.
What You'll Do
The Service Assurance Systems (SAS) Group, part of Global Operations, develop and maintain many software systems and applications which support the operation of Viasat services.
The team collects assurance and operational data from a wide range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose, the team builds and maintains data pipelines that cleanse, transform and enrich raw data, making it readily consumable by our stakeholders across operations, engineering and management.
As a Site Reliability Engineer (SRE), you will play a key role in bridging the gap between software development and operational reliability. You will be responsible for supporting the applications built and maintained by the SAS group, managing deployments, investigating operational issues, and driving the team’s transition toward a modern DevOps culture. You will act as a first point of contact for service issues raised through the ServiceNow ticketing system, working to diagnose, triage and resolve incidents efficiently while collaborating closely with developers and operations teams.
The day-to-day
You will be working as part of a small team of developers and operations engineers, supporting the evolution of Viasat’s network and service monitoring capabilities, ensuring it remains world class in support of existing and future services.
Day-to-day the role will involve:
- Investigate and resolve incidents, service requests and problems raised through the ServiceNow ticketing system, ensuring timely and thorough resolution within agreed SLAs.
- Manage and oversee application deployments across development, staging and production environments, ensuring smooth and reliable release processes.
- Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation.
- Own and maintain CI/CD pipelines, working to improve build, test and deployment automation to reduce manual effort and increase release confidence.
- Champion and drive the adoption of DevOps practices and culture within the SAS group, working towards greater automation, infrastructure-as-code, and operational maturity.
- Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure.
- Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency.
- Collaborate with software developers to identify recurring operational issues and feed findings back into the development process to improve application resilience.
- Maintain clear and up-to-date operational documentation including runbooks, deployment guides and incident post-mortems.
- Provide written and verbal progress updates on open incidents, deployments and operational improvements to the SAS group and wider stakeholders.
- Support on-call and out-of-hours incident response as required.
- Liaise with engineering and infrastructure teams to ensure system changes are communicated and operationally risk-assessed before deployment.
- Familiarity with both Windows and Linux operating systems, including command-line administration and troubleshooting.
- Demonstrable experience with SQL or data pipeline technologies.
- Experience with data processing frameworks such as Apache Airflow, Apache Flink or Google Dataflow, with a practical understanding of ETL pipeline design and the ability to diagnose issues across data transformation and scheduling workflows.
- Demonstrable CI/CD experience — building, maintaining and improving pipelines using tools such as Jenkins, GitLab CI, GitHub Actions or equivalent.
- Hands-on experience with AWS services (e.g. EC2, S3, ECS, Lambda, CloudWatch) and containerisation using Docker.
- Experience with monitoring and observability tooling, including Prometheus and AWS CloudWatch, for metrics collection, alerting and dashboarding.
- Proficiency in scripting with Python for automation, tooling and operational support tasks.
- Strong communication skills — able to clearly articulate technical issues and their impact to both technical and non-technical audiences, and to provide timely updates on incident progress.
- Experience using an ITSM/ticketing system (e.g. ServiceNow) for incident management, request fulfilment and problem tracking.
- A proactive, solution-oriented approach with strong attention to detail and a commitment to operational excellence.
- A reasonable understanding and appreciation of IT and security standard methodologies in an operational environment.
- Understanding of TCP/IP networking principles, including the ability to diagnose connectivity issues and interpret network traffic.
- Familiarity with event streaming platforms such as Apache Kafka, including an understanding of producer/consumer patterns and how message queues are used to move data reliably between systems at scale.
- Familiarity with infrastructure-as-code tools such as Terraform or Ansible.
- Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights.
- Exposure to Kubernetes or other container orchestration platforms.
- Experience working in an Agile or DevOps team environment.
Viasat is proud to be an equal opportunity employer, seeking to create a welcoming and diverse environment. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, ancestry, physical or mental disability, medical condition, marital status, genetics, age, or veteran status or any other applicable legally protected status or characteristic. If you would like to request an accommodation on the basis of disability for completing this on-line application, please click here.
Cara melamar
Untuk melamar pekerjaan ini, Anda perlu otorisasi di situs web kami. Jika Anda belum memiliki akun, silakan daftar.
Posting CVPekerjaan serupa
Senior Hook Up & Commission Tech (Instrument)
McDermott International, Ltd,
Batam, Riau
1 minggu yang lalu
Job Overview JOB DESCRIPTION The Senior Hook Up & Commission Tech completes a variety of Hook Up & Commissioning assignments as needed and can complete work with a limited degree of supervision. They are an informal resource for colleagues with less Hook Up & Commissioning experience. The Senior Hook Up & Commission Tech has developed proficiency in a range of...
Territory Manager Batam
Home Credit Indonesia,
Batam, Riau
1 minggu yang lalu
About This JobHome Credit IndonesiaLocation: Batam, Riau Islands, IndonesiaWork Mode: On-siteIndustry: Appliances, Electrical, and Electronics ManufacturingJob DescriptionAs a Territory Manager, you'll have the opportunity to lead and inspire a dynamic team, drive business growth, and shape the future of our organization in your region.If you're interested in exploring this exciting opportunity, reach out to us today to learn more or...
Manager Construction (Fab)
McDermott International, Ltd,
Batam, Riau
2 minggu yang lalu
Job DescriptionJob Overview:The Manager Construction (Fab) role requires an in-depth understanding of Construction concepts, theories, and principles and basic knowledge of other related disciplines. The Manager Construction (Fab) must be able to apply an understanding of the industry to improve effectiveness, provide guidance, and influence processes and policies for the Construction discipline as well as identify and resolve technical, operational,...