Engagement Manager - Support Lead in | ITJOBCELL

ENGAGEMENT MANAGER - SUPPORT LEAD
  • 0 Applied
  • 3
  • 0 - 0 Years
  • Not disclosed
  • Engagement Manager - Support Lead
  • Today
Job Description
What you’ll do: As an Engagement Manager and Techno-Functional Support Lead for the Production Engineering (PE) and Service Delivery (SD) stream, you will be the primary interface between our critical healthcare AI platforms and our key stakeholders, both internal and external. You will own the strategic direction, reliability, scalability, and availability of these platforms, moving beyond day-to-day ticket resolution to define and implement a comprehensive Support Strategy rooted in Site Reliability Engineering (SRE) principles. You will lead a high-performing team of engineers, guiding them through complex incident resolutions, orchestrating advanced cloud automation, and expertly bridging the gap between Data Engineering, ML Engineering, Operations, and our business stakeholders. This role demands a unique blend of deep technical acumen, strategic thinking, and exceptional client-facing communication skills. You will serve as the primary escalation point, technical architect, and relationship manager for support operations, ensuring high availability and optimal performance for multi-geography projects while aligning technical solutions with business objectives. Key Responsibilities: Strategic Engagement & Leadership: Client & Stakeholder Management: Serve as the primary technical and functional point of contact for key clients and internal business units regarding platform reliability, performance, and service delivery. Translate complex technical issues and resolutions into clear, business-centric communications. Support Strategy & Roadmap: Define, champion, and execute the long-term support strategy, incorporating SRE principles (SLOs, SLIs, Error Budgets) to proactively enhance platform reliability and user experience. Align this strategy with overall product and business goals. Team Leadership & Development: Lead, mentor, and empower a team of Senior and Junior Support Engineers. Foster a culture of continuous improvement, technical excellence, and client-centricity. Conduct performance reviews, manage shift rosters, and facilitate knowledge sharing. Major Incident Management & Communication: Act as the Major Incident Manager and the primary escalation POC. Lead technical resolution war rooms, drive Root Cause Analysis (RCA), and own the Corrective Action processes. Crucially, manage all stakeholder communications during critical incidents, providing timely and transparent updates. Technical Architecture & Automation: SRE Implementation & Governance: Drive the adoption and continuous improvement of SRE practices, defining and tracking key metrics (SLOs, SLIs, Error Budgets) to ensure platform reliability and performance meet business expectations. Infrastructure as Code (IaC) & Automation: Architect and enforce IaC best practices using Terraform/CloudFormation, ensuring environments are self-healing, reproducible, and compliant. Oversee the design and implementation of comprehensive monitoring, logging, and alerting stacks to enable proactive support. Release & Deployment Management: Oversee the maturity of the release pipeline, ensuring zero-downtime deployments, automated rollbacks, and robust change management processes. Cloud-Native Innovation: Introduce and integrate advanced cloud-native tools and services to improve operational efficiency, enhance automation, and drive cost optimization. Collaboration & Influence: Design for Supportability: Collaborate closely with Solution Architects, Product Managers, and Engineering teams to influence the design of new features and platforms, ensuring they are inherently supportable, maintainable, and scalable from inception. Cross-Functional Alignment: Bridge the gap between Data Engineering, ML Engineering, and Operations, ensuring seamless integration and efficient resolution of cross-functional issues. Process Optimization: Continuously evaluate and optimize support processes, leveraging automation and best practices to improve efficiency and reduce Mean Time To Resolution (MTTR).
Skills
Expert Proficiency in Cloud Infrastructure (AWS/GCP preferred): Compute & Serverless: Deep expertise in configuring and optimizing EC2/GCE instances, Lambda/Cloud Functions, and auto-scaling groups based on custom metrics. Networking: Advanced knowledge of VPC peering, Transit Gateways, Load Balancers (ALB/NLB), Route53/Cloud DNS, and VPN/Direct Connect setups. Security & IAM: Ability to design least-privilege IAM policies, manage KMS keys for encryption, and implement WAF rules to protect public-facing endpoints. Storage Lifecycle: Managing tiered storage strategies (S3/GCS lifecycle policies) and high-performance block storage optimization (EBS/Persistent Disks). Software Engineering Proficiency: Core Development: Strong command of Python (using frameworks like Flask/Django/FastAPI) or Java (Spring Boot) to build production-grade applications and tooling. White-Box Debugging: Ability to clone application repositories, navigate complex codebases, attach remote debuggers, and identify the exact line of code causing logic errors or memory leaks. Code Contribution & Review: Comfortable submitting Pull Requests (PRs) for bug fixes, conducting thorough code reviews for peers, and writing robust unit/integration tests to prevent regressions. Performance Profiling: Experience using profilers (e.g., cProfile, JProfiler) to analyze thread dumps and heap dumps to resolve performance bottlenecks. Microservices & API Architecture: Protocols & Patterns: Deep understanding of RESTful API design, gRPC protobufs, and asynchronous communication patterns (Pub/Sub, Kafka, SQS). Resiliency Patterns: Implementation of Circuit Breakers, Retry logic (exponential backoff), and Rate Limiting to prevent cascading failures. Service Mesh: Hands-on experience with traffic splitting, mutual TLS (mTLS), and observability sidecars. Distributed Tracing: Analyzing request flows across microservices using tools like OpenTelemetry to pinpoint latency hotspots. Container Orchestration (Kubernetes - GKE/EKS preferred): Cluster Administration: Managing GKE/EKS upgrades, node pools, and understanding the control plane components (API Server, Scheduler, Controller Manager, etcd). Networking & Security: Configuring Ingress Controllers (Nginx/ALB), Network Policies (Calico/Cilium), and Pod Security Policies/Contexts. Resource Management: Tuning resource requests/limits, setting up Horizontal/Vertical Pod Autoscalers (HPA/VPA), and debugging errors. Helm & Operators: Writing and managing complex Helm charts and understanding how Kubernetes Operators automate stateful applications. Automation & Tooling: Internal Tools Development: Building custom CLI tools or web dashboards to empower L1/L2 support teams (e.g., a "one-click" user reset tool or a log analyzer dashboard). Event-Driven Automation: Creating self-healing workflows (e.g., a Lambda function that automatically restarts a hung process or cleans up temporary files when disk alerts fire). Infrastructure as Code (IaC): Terraform Mastery: Structuring Terraform projects using reusable modules, managing remote state locking (S3 + DynamoDB), and using workspaces for multi-environment management. Configuration Management: Using Ansible playbooks for OS-level hardening, patch management, and configuration drift detection. Policy as Code: Implementing tools like OPA (Open Policy Agent) or Sentinel to enforce compliance checks within the IaC pipeline (e.g., preventing public S3 buckets). Databases (SQL & NoSQL): Relational (PostgreSQL): Analyzing outputs to optimize slow queries, managing connection pooling (PgBouncer), and configuring replication/WAL archiving for disaster recovery. NoSQL (Cassandra/MongoDB): Understanding consistency levels, partition keys, sharding strategies, and diagnosing compaction or garbage collection issues. Caching: Implementing and troubleshooting caching layers (Redis/Memcached) to offload database read pressure. DevOps Toolchain: CI/CD Pipelines: Designing complex Jenkins pipelines (Groovy shared libraries) that include parallel build stages, automated testing, and canary deployments. Version Control: Advanced Git strategies (Gitflow/Trunk-based), handling large merge conflicts, and using git hooks for pre-commit checks. Artifact Management: Managing Docker registries and dependency repositories (Artifactory/Nexus) with lifecycle policies to clean up old builds.

Suggested Jobs

Gruppo Zenit India Pvt Ltd, subsidiary of Gruppo Zenit Srl, Italy, with more than 30 years of proven success and industry experience, specializes in delivering cutting edge Software Solutions and IT Services to a predominantly European client base.We at Gruppo Zenit accompany organizations on their Digital Transformation journey, offering the required technical expertise and delivering customized solutions that foster sustainable growth through technological innovation. At Gruppo Zenit, we collaborate closely with our clients to generate measurable results through digital and infrastructure transformation. Our expertise covers the design and development of web and mobile applications, ERP integration solutions, and IT infrastructure management. We ensure seamless project execution and provide continuous support throughout the entire application life cycle, including corrective and evolutionary maintenance services. Position Overview We are looking for a highly skilled Oracle Cloud Infrastructure (OCI) Architect to join our team and support strategic enterprise projects. In this role, you will be responsible for managing, migrating, and optimizing complex infrastructures and applications across cloud and hybrid environments. You will play a key role in defining scalable, secure, and high-performing architecture based on Oracle Cloud Infrastructure, working closely with cross-functional teams and client stakeholders. You will also act as a trusted advisor to clients, ensuring that cloud solutions are aligned with business objectives, technical requirements, and best practices. Key Responsibilities: Analyze technical and business requirements to design cloud architecture based on Oracle Cloud Infrastructure Design, implement, and manage OCI cloud infrastructures, integrating IaaS, PaaS, and native OCI services Collaborate with development, operations, and security teams to optimize workflows and DevOps processes Monitor performance, security, and compliance of OCI environments, applying cloud best practices Provide technical consulting to clients and support operational teams Manage cloud infrastructure lifecycle, including migration and optimization activities Ensure service reliability and compliance with SLA requirements Experience & Qualifications: Master’s degree in computer science, Engineering, or a related field (or equivalent experience) Proven experience as a Cloud Architect in enterprise environments Strong knowledge of Oracle Cloud Infrastructure (OCI) Experience in designing and managing cloud and hybrid infrastructures Experience with cloud monitoring and performance optimization Experience in operational coordination and incident management in line with SLA requirements

Experience Level - Specialist Years of Experience - 7-8 years Reports To - SW Engineering Manager Employment Type - Full-time Location - HEX20LABS PVT LTD, Trivandrum Job Code - HEX20_SW_J01003 Role Summary :- We are looking for a Senior Embedded Software Engineer to architect and lead develo pment of real-time and Linux-based embedded systems. You will drive technical decisions across FreeRTOS and embedded Linux platforms, set architecture and coding standards and mentor engineers across the team. Key Responsibilities :- Architect embedded software systems spanning FreeRTOS-based subsystems and embedded Linux platforms Define system architecture, task/process partitioning, IPC strategy, and overall software design Lead technical decisions on RTOS vs. Linux boundaries, board support packages (BSPs), and kernel/driver integration Set coding standards, review processes, and best practices across the embedded team Guide and mentor junior and mid-level engineers; lead design and code reviews Drive root-cause analysis for the most complex system-level and cross-layer issues Own technical roadmap decisions in collaboration with product and hardware teams Evaluate and integrate new SoCs, toolchains, and platform technologies Required Skills & Qualifications :- 5-7 years of embedded software engineering experience, including system-level architecture ownership Proven experience architecting solutions using FreeRTOS (task/priority design, memory strategy, ISR-to-task communication, scheduling trade-offs) Strong hands-on experience with embedded Linux (kernel configuration, device drivers, BSP customization, Yocto/Buildroot) Deep understanding of microcontroller/microprocessor architectures (ARM Cortex-M/A or similar) Strong grasp of low-level debugging and system bring-up across both RTOS and Linux environments Track record of leading design reviews and mentoring engineers Excellent communication skills, able to justify architectural trade-offs to cross-functional stakeholders Ability to lead, mentor, and do task planning for a group size varying from 3 to max 7 engineers.

SRS Global Technologies Pvt. Ltd. is a leading software development company specializing in innovative technology solutions for the U.S. healthcare industry. As the dedicated product development center for SRS Web Solutions (USA), we have been delivering cutting-edge digital healthcare solutions since 2014. Our flagship Clinic Management System is a comprehensive, paperless platform designed to simplify and enhance practice management for Dentists, Physicians, Veterinarians, and Optometrists. In addition, our Home Healthcare Automation System streamlines operations for Home Healthcare Agencies and Private Duty Nursing providers, enabling greater efficiency and improved patient care. Beyond healthcare software, SRS Global offers end-to-end technology services, including Web Development, Mobile Application Development, and Digital Marketing. Our experienced team collaborates closely with clients throughout the entire product lifecycle—from concept and strategy to development, deployment, and ongoing support—ensuring high-quality, scalable, and innovative solutions. Key Responsibilities Product Positioning & Go-to-Market Develop product positioning, messaging, and value propositions for mConsent solutions. Create customer-centric messaging for dental practices, DSOs, orthodontists, pediatric practices, oral surgeons, and other specialties. Plan and execute go-to-market strategies for new products, features, and enhancements, including launch campaigns, customer communications, webinars, training, and sales enablement. Sales Enablement & Competitive Intelligence Create and maintain sales assets including presentations, one-pagers, battle cards, competitive comparison sheets, ROI calculators, demo scripts, FAQs, and customer success stories. Conduct product training for Sales and Customer Success teams. Monitor competitors, industry trends, and market opportunities while delivering actionable insights and win/loss analysis. Content, SEO & AI Search Optimization Develop product-focused content including landing pages, solution pages, blogs, case studies, webinars, whitepapers, videos, and customer success stories. Drive SEO, AEO (Answer Engine Optimization), and GEO (Generative Engine Optimization) initiatives to improve visibility across search engines and AI-powered platforms. Partner with development and content teams to optimize technical SEO, schema markup, website performance, Core Web Vitals, and search rankings. Customer Marketing & Demand Generation Support customer advocacy programs, testimonials, review campaigns, webinars, newsletters, and expansion initiatives. Collaborate with Demand Generation teams on paid advertising, email marketing, partner marketing, conversion optimization, and campaign messaging to drive pipeline growth. Analytics & Performance Measure and report on product adoption, feature usage, campaign performance, website traffic, AI search visibility, pipeline influence, win/loss trends, customer engagement, and marketing-attributed revenue. Provide data-driven recommendations to improve product adoption, customer engagement, and marketing ROI.