Hadoop Administrator & Data Engineer in Trivandrum | ITJO...

HADOOP ADMINISTRATOR & DATA ENGINEER
  • 0 Applied
  • 4
  • Trivandrum
  • 4 - 0 Years
  • Not disclosed
  • Hadoop Administrator & Data Engineer
  • Today
Job Description
Operate and tune HDFS, YARN, Hive, and Spark across both clusters, including a mixed-version estate: one cluster runs an older HDP-era stack with Spark 2.x and a legacy Python driver environment; the other has recently been upgraded to a modern Hive release. Managing that version skew — and the compatibility traps it creates for job authors — is part of the role. Own gateway host capacity. Monitor and set limits on concurrent worker pools, multiprocessing fan-outs, and Spark drivers so a single workload cannot starve the box. Push back on resource requests that the available evidence doesn't actually support. Run the Hive Metastore and HiveServer2, including JVM sizing and OOM exposure. Diagnose transient failures such as partition-add serialization errors and thrift session rejections. Administer the clusters through Ambari: service configs, alerts, rolling restarts, and version upgrades. Handle capacity planning, HDFS quotas, and data retention. Every dataset we generate needs a documented retention and cleanup story, and you'll help enforce that on the storage side. Partner with the data engineering team, which orchestrates pipelines with Apache Airflow on AWS MWAA. Most remote work reaches the clusters over SSH from the gateways, so you'll be the escalation point when a pipeline task fails for infrastructure reasons rather than logic reasons. Manage access control, SSH key hygiene, and service accounts across gateways and cluster nodes. Build the monitoring and alerting we're missing — host load, per-tenant resource consumption, metastore health, and early warning before a gateway becomes unresponsive.
Skills
Should work in US hours ( 6am Pacific Time to 5 Pacific Time ) 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent). Deep, hands-on HDFS, YARN, and Hive administration — not just usage. You should be comfortable reading NameNode and metastore logs and reasoning about JVM heap behavior. Strong Linux systems administration: process and memory forensics, SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts. Spark-on-YARN operational experience, including memory overhead tuning and diagnosing driver-side failures. Solid scripting in Bash and Python. Judgment about shared infrastructure. We want someone who will say "that benchmark measured database throughput, not gateway capacity — start lower and ramp" instead of maximizing a number. Ambari administration and in-place Hadoop version upgrades. Apache Airflow, particularly AWS MWAA. MySQL and MongoDB operations at scale. Elasticsearch or Sphinx/Manticore search infrastructure. Terraform and AWS (S3, EMR, IAM). Experience decommissioning or migrating legacy Hadoop estates onto newer platforms. Send your resume to careers@polussolutions.com Please include “Hadoop Administrator & Data Engineer - [Your Name]” in the subject line

Suggested Jobs

Assist in the development of edit specifications, based on any available global medical standards, therapeutic area standards, and the protocol, used to clean the study.? Performs user acceptance testing (UAT) on eCRF build and edit specifications.? Creates supporting DM process documentation to LDM and/or performs peer review of documentation, including updating documentation.? Support the coding schedule defined in the data management plan.?Collaborate with data coding specialists on a regular basis to guarantee timely coding.? Supports/maintains quarterly coding review cycles.? Performs manual data listing reviews and submits queries as?appropriate. Assist with and/or performs user acceptance testing of lab data?standards. Evaluates quality of lab data entry and addresses inconsistencies with sites and CRAs as applicable.? Assists in the SAE reconciliation process. This may include coordination with medical experts and Global Drug Safety.?Follow up on discrepancies and resolve so both databases are consistent.? Applies criteria for subject stage gate of No More Issues (NMI).?Also, must coordinate and review medical and statistical queries and certify they are adequately resolved.? Assist in the development of a blind review report and conducting a blind review meeting to assign patient validity.? Assist in developing and generating study report listings according to ICH and if present company guidelines.? Coordinate the query management system functions.? Perform the final patient review and database lock activities. Assist in coordinating the processing of scheduled data transfers (PK/PD data, imaging data, Laboratory data) from external vendors and performs relevant review/reconciliation.? Review query responses and ensure data quality.? Reviews Site responses to queries and evaluates the necessity of a re-query. If applicable, communications with Site Coordinators are performed for resolution.? Attends and may lead internal and external team meetings.? Reviews and/or provides meeting minutes.? Supports training and development of Clinical Data Coordinators. Assists with eCRF design. May be required to develop the eCRF and/or provide peer review. May serve as a back up to the LDM for internal and external study teams.

Duration: 3 Months Mode: online & offline Stipend: Basic stipend provided Certificate: Internship Certificate on successful completion Who Can Apply IGNOSI is looking for enthusiastic interns who are willing to learn and contribute. Intern Responsibilities Content development for sales tools Weekly reporting and status updates Data updating for an AI face-detection application Data management and documentation Multi-language content support

Accurately, within project timelines and according to project guidelines, codes clinical trials data (e.g., adverse events, medical history, physical examination findings, and/or medications) and ensures completeness, accuracy, and consistency of the codes through use of data listings, computer generated reports, or on-line review. Develops and validates coding specifications, ad hoc listings, reports, and queries for use in the validation of coded data in the clinical database. Performs study specific User Acceptance Testing (UAT) of the coding module used for assigned studies. Performs peer reviews of coding reports and provides guidance to other Data Coding Specialists. Coordinate up versioning of assigned dictionaries throughout the life of a study, as needed. Review or provide input into the development of SOPs, Data Management Plans, and project management tools. Maintains a good understanding of currently used coding dictionaries and guidelines in order to effectively serve as a department resource. Interfaces with project teams to resolve problems and issues dealing with the coding of clinical data. Trains Data Coding Specialists, as well as Data Management staff, on coding dictionary structures, conventions and/or use of coding software to apply codes or run reports. Interact with Sponsors by gaining working knowledge of project coding specifications, receiving feedback on coding reports, promoting consistency across projects within or resolving coding questions. Verifies active sponsor dictionary subscriptions.