Job Description
Role: Data Bricks Migration and Support engineer
\n
Location: Seattle, WA / Bellevue, WA / Everett, WA / Renton, WA / Richardson, TX / Plano, TX / Dallas, TX / St. Louis, MO / Charleston, SC / Arlington, VA
\n
\n
Responsibilities
\n
Support post-migration environment from IBM DataStage to Databricks
\n
Incident & Lifecycle Management
\n
- \n
- CI/CD Deployment: Support code deployments across Development, Test, and Production environments using Databricks Repos and REST APIs
- Monitoring & Alerting: Set up monitoring via Databricks System Tables and observability tools to catch job failures, data anomalies, or latency spikes early
\n
\n
\n
Pipeline Maintenance & Orchestration
\n
- \n
- Workflow Management: Transition from DataStage job sequences to native data bricks workflows for scheduling, dependency tracking, and alerts
- ETL Refactoring: Troubleshoot and fix issues in generated PySpark or Spark SQL code that replaced legacy DataStage Transformer or Lookup stages
- Streaming & Batch Integration: Support ongoing data ingestion using data bricks autoloader to process files continuously from cloud storage
\n
\n
\n
\n
Performance Tuning & Cost Optimization
\n
- \n
- Compute Management: Monitor and configure serverless or classic clusters to prevent over-provisioning
- Query Optimization: Analyze Spark execution plans. Replace inefficient row-by-row processing logic (a common DataStage carryover) with vectorized operations and native Spark functions
- Storage Optimization: Maintain Delta Lake tables by enforcing layout optimization (\(ZORDER\)
\n
\n
\n
\n
Data Governance & Security
\n
- \n
- Access Control: Implement granular permissions, column-masking, and row-level filters using Data bricks unity catalog to replace DataStage's legacy security policies
- Data Quality: Utilize Delta Live Tables (DLT) to build pipelines with built-in, declarative data quality expectations and monitoring
\n
\n
\n
