About the role
This is a remote position.
Project Description
* Support secure, reliable, and compliant operation of cloud-hosted research applications, agentic AI systems, scientific SaaS platforms, and production services. * Work across GCP, AWS, cloud containers, AI/ML models, MCP servers, APIs, scientific databases, and third-party SaaS platforms. * Support scientists, developers, IT, cybersecurity, networking, data teams, architecture, and vendors in transitioning research prototypes into secure, repeatable, production-ready services.
Key Responsibilities
* Administer GCP/AWS cloud projects, IAM, service accounts, networking, storage, compute, quotas, and SaaS environments. * Operate Cloud Run, Docker containers, registries, DNS, TLS, load balancing, and Google Cloud IAP. * Manage identities, RBAC, SSO, OAuth, service accounts, machine-to-machine authentication, secrets, API keys, tokens, and certificates. * Maintain approved AI agents, MCP servers, tools, integrations, data sources, permissions, and credentials. * Maintain Dockerfiles, container images, Python/application dependencies, registries, runtime configurations, and rollback procedures. * Monitor logs, metrics, traces, alerts, dashboards, health checks, API limits, model/tool failures, token usage, cloud costs, and service availability. * Maintain CI/CD pipelines and infrastructure-as-code, including source control, security scanning, testing, approvals, and deployment traceability. * Perform incident triage, escalation, root-cause analysis, backup/recovery, disaster recovery, and service continuity. * Manage incidents, service requests, problems, and changes using ServiceNow and Jira. * Support vulnerability remediation, patching, security investigations, audits, threat modeling, and risk assessments. * Identify shadow AI services, unmanaged integrations, unapproved MCP servers, and overprivileged identities. * Create operational runbooks, technical documentation, and lifecycle/ownership documentation. * Help transition scientist-managed prototypes into secure, supportable production deployments.
Requirements
Mandatory Requirements
* Bachelor’s degree in Computer Science, Information Systems, Engineering, Technology, or related field. * 3+ years of experiencein systems/cloud administration. * Production administration experience with GCP, AWS, Azure, or comparable cloud platforms. * Hands-on experience with: * Linux * Docker/containers * Python environments * APIs * Networking, DNS, and TLS * IAM, SSO, OAuth, service accounts * Secrets management * Serverless/container orchestration such as Cloud Run or Kubernetes * Knowledge of: * CI/CD * Infrastructure as Code * Monitoring and observability * Incident management * Backup/recovery * Patching * Vulnerability remediation * Experience with ServiceNow and Jira. * Strong troubleshooting, documentation, communication, and cross-functional collaboration skills.
Preferred Skills
* GCP IAP, Cloud Run, Secret Manager, Artifact Registry, Cloud Monitoring. * Vertex AI, Gemini, or other AI/ML platform administration. * Experience supporting GCP + AWS hybrid/multi-cloud environments. * Agentic AI, AI orchestration, MCP servers, AI APIs, or model operations. * Scientific computing, bioinformatics, laboratory automation, multiomics, research data platforms, or scientific SaaS. * Container scanning, SBOM, software supply-chain security, data integrity, and auditability. * Cloud, systems engineering, ITSM, or cybersecurity certifications.
Originally posted on Himalayas