عن الوظيفة
About the RoleWithin Earth is looking for a highly experienced Senior Infrastructure Engineer to own and manage our mission-critical production infrastructure.Our hotel platform processes over 1 billion searches per day and thousands of bookings daily.
This role is responsible for ensuring our production environment remains stable, secure, scalable, and available 24/7 with near-zero downtime.You will be responsible for managing our virtualization platform, Linux and Windows production servers, enterprise databases, load balancing, monitoring, automation, performance optimization, and troubleshooting across the full infrastructure stack.We are looking for someone who can work independently, make sound technical decisions, quickly identify bottlenecks, perform Root Cause Analysis (RCA), and continuously improve the platform.What You'll OwnProduction InfrastructureXCP-ng Virtualization ClusterLinux & Windows Production VMsHAProxy ClusterSQL ServerPostgreSQLMongoDBRedisClickHouseKubernetes & DockerMonitoring & Alerting PlatformGitHub CI/CDInfrastructure AutomationAzure InfrastructureProduction Performance & Capacity PlanningRequired Technical SkillsInfrastructure & VirtualizationYou should be able to independently:Deploy, manage, troubleshoot and optimize XCP-ng clustersManage VM lifecycle, storage, networking and live migrationDiagnose virtualization performance bottlenecksPlan infrastructure capacity and resource allocationPerform disaster recovery and backup managementLinux & Windows Production ServersYou should be able to independently:Deploy and maintain production Linux and Windows serversTroubleshoot CPU, memory, storage and networking issuesPerform OS performance tuningHandle production incidentsSecure and harden production serversAutomate administration tasksHAProxyAdvanced production experience with:HAProxy deployment and administrationHigh Availability clustersHealth ChecksSSL TerminationSticky SessionsACL RulesBackend FailoverLoad BalancingZero Downtime deploymentsPerformance tuningTroubleshooting production issuesDatabase AdministrationHands-on production administration of:SQL ServerPostgreSQLMongoDBRedisClickHouseYou should be able to independently:Install and configure databasesBackup & RestoreReplicationHigh AvailabilityPerformance tuningQuery optimizationCapacity planningDatabase monitoringTroubleshoot production issuesRoot Cause AnalysisMonitoring & ObservabilityExperience building and managing enterprise monitoring using:ZabbixDatadogGrafanaPrometheusELK (preferred)You should be able to:Build dashboardsConfigure alertsMonitor infrastructure healthMonitor database performanceDetect anomaliesInvestigate incidentsIdentify performance bottlenecksKubernetes & DockerProduction experience with:KubernetesCluster AdministrationDeploymentsServicesIngressHelmStatefulSetsPersistent VolumesAutoscalingRolling UpdatesProduction TroubleshootingDockerDockerDocker ComposeImage OptimizationContainer NetworkingPrivate RegistryProduction OperationsCI/CD & AutomationExperience with:GitHubGitHub ActionsCI/CD PipelinesRelease AutomationRollback StrategySecrets ManagementStrong scripting skills using:PowerShellBashPythonAutomation experience is essential.CloudExperience managing Azure production environments including:Virtual MachinesNetworkingStorageMonitoringHybrid InfrastructurePerformance & TroubleshootingYou should be confident in:Finding production bottlenecksPerformance optimizationCapacity planningInfrastructure scalingDatabase performance tuningApplication infrastructure troubleshootingRoot Cause Analysis (RCA)High Availability designZero Downtime operationsWe're Looking For Someone Who Can✅ Own production infrastructure independently.✅ Troubleshoot across the complete infrastructure stack.✅ Identify problems before they impact production.✅ Make sound technical decisions.✅ Optimize systems before recommending hardware upgrades.✅ Automate repetitive operational tasks.✅ Manage high-availability infrastructure.✅ Design scalable production systems.✅ Work comfortably in a fast-moving 24/7 production environment.
7+ years in Senior Infrastructure / Platform Engineering.Experience supporting enterprise production platforms.Experience managing high-availability environments.Strong production troubleshooting and RCA skills.Experience with both on-premises infrastructure and Microsoft Azure.
7+ years in Senior Infrastructure / Platform Engineering.Experience supporting enterprise production platforms.Experience managing high-availability environments.Strong production troubleshooting and RCA skills.Experience with both on-premises infrastructure and Microsoft Azure.