We are seeking a professional to design and build back-end services that support our portfolio of data-centric clinical and analytic applications. These applications leverage
cloud computing, big data,
mobile, data science, data warehousing, machine learning using state-of-the-art software development applications and frameworks.
- Ensure that cloud-based micro-services adhere to uptime and accuracy targets, are resilient, and scale as data volumes and traffic increase.
- Work closely with the data engineering, platform, and solutions teams to develop applications as required to benefit our practice and patients.
- Collaborate with Product Owners, Product Managers, Architects to translate requirements into code.
- Develop services around data warehousing, big data, cloud computing, business intelligence, analytics, and machine learning.
- Participate in DevOps, Agile, continuous development, and integration frameworks.
- Program in high-level languages such as Go, Python, Java, etc.
- Ensure all appropriate documentation of processes and source code is created and maintained.
- Communicate effectively with peers, leaders, and customers throughout the organization.
- Participate in expert-level troubleshooting and resolve problems through root cause analysis, data, and system investigation.
- Contribute to design and architecture discussions with Principals and Architects.
- Lead targeted cross-functional improvement efforts and mentor more junior software engineers.
- Solve complex problems; take a new perspective on existing solutions.
- Work independently with minimal guidance. You may lead projects or project steps within a broader project or have accountability for ongoing activities or objectives.
- Act as a resource for colleagues with less experience.
Data Engineering Skills & Experience
- Create, verify, and maintain data replication scripts.
- Create, verify, and maintain data validation, processing, and ingestion pipelines.
- Deploy and automate the execution of data replication scripts and data pipelines in cloud infrastructure.
- Create and maintain data catalogs that describe datasets and their contents (i.e., files, file types, tables/views, columns, fields, etc.).
- Create, verify, and maintain dashboards and reports that characterize ingested datasets.
- Create, verify, and maintain data validation scripts/APIs that verify the production dataset contains the correct number of samples/records, expects values/fields/columns are populated, and values are of the correct data type, format, and range.
- Deploy and automate the execution of data validation scripts/APIs.
- Create and maintain user documentation (dataset descriptions, tutorials, code examples, etc.).
- Define entitlements, user groups, roles, and permissions utilized to grant access to datasets.
Programming Languages
- Primary pipeline development language will be Python.
- Some datatypes and formats may require the use of other languages (i.e., Java, R, etc.) because the libraries/frameworks/sdks available to work with those datatypes and formats are not available in Python.
Operating Systems
- Primary operating system for data pipeline execution will be Linux, with data pipelines packaged, deployed, and run as containers.
- Data source systems could be Windows or Linux based.
Infrastructure
- Primary data platform and data pipeline execution infrastructure will be hosted on Google Cloud Platform (GCP) utilizing cloud-native technologies (i.e., Google Cloud Storage, BigQuery, Google Batch, Dataflow, Cloud SQL, etc.).
- Data will be replicated from various on-premises sources that include laboratory instruments, network shared drives, and Windows desktops attached to instruments.
Development Tools
- Sprints, features, and tasks will be managed in Azure DevOps.
- Code will be managed and versioned in Azure DevOps-based git repositories.
- Code will be compiled, packaged, and deployed utilizing Azure DevOps build pipelines.
- Data pipelines will be packaged, deployed, and run in Docker containers.
- Docker containers will be stored and versioned in Google Cloud Artifact Repositories.
- Veracode will be utilized to scan source code for vulnerabilities and Prisma Cloud will be utilized to scan containers.
- The standard integrated development environment will be JetBrains (PyCharm, IntelliJ, etc.) or VSCode.
Preferred Candidates
- Experience working on healthcare, life science, or scientific research projects.
- A degree or domain knowledge in a life science-related field (biochemistry, genetics, biology, etc.).
- Experience with Google Cloud Platform-based infrastructure and services.
- 100% remote.