GATX is a global company with a collaborative culture and a focus on providing transportation and data-driven business solutions. The Data Engineer will help evolve GATX’s data platform into a scalable, self-service data and AI ecosystem by developing and supporting data integration, ETL, cloud-based data pipelines, and analytics platform components. The role also supports platform operations, production issue resolution, continuous improvement, and innovation across AWS and Databricks environments.
Participate in all phases of the Software Development Life Cycle (SDLC) for data integration, ETL, and cloud-based data pipeline development, including requirements analysis, design, development, testing, deployment, and production support
Assist in designing, developing, testing, deploying, and supporting batch and near-real-time data pipelines on AWS and Databricks
Develop data transformations using SQL, Python, Spark, Delta Lake, and established data engineering patterns
Participate in requirements analysis, estimation, solution design, code reviews, testing, and production deployment
Develop and maintain Databricks notebooks, workflows, jobs, Delta tables, and related data components
Follow established architecture, coding, DevOps, security, governance, and documentation standards
Create unit tests and assist with integration, reconciliation, performance, and regression testing
Investigate data quality, pipeline, and production-processing issues, escalating complex problems when appropriate
Create and maintain source-to-target mappings, transformation rules, data flow diagrams, and operational documentation
Collaborate with senior engineers, architects, analytics teams, business stakeholders, infrastructure teams, and external partners
Build an understanding of supported business processes, source systems, reporting requirements, and upstream and downstream dependencies
Participate in platform upgrades, deployment automation, monitoring improvements, and operational process enhancements
Continuously develop expertise in AWS, Databricks, Spark, Python, SQL, Delta Lake, Unity Catalog, and modern data engineering practices
Assist in monitoring the health, performance, and reliability of data pipelines, ETL processes, and cloud-based data platform components
Support established monitoring, alerting, and operational procedures to help ensure data processing service levels and platform availability are maintained
Monitor daily production processes and work with IT administrators, support teams, and other DARI team members to investigate, troubleshoot, and resolve production issues in a timely manner
Maintain an understanding of established service level agreements (SLAs) and provide support that meets operational and business expectations
Participate in continuous improvement initiatives focused on reducing production incidents, improving system stability, and enhancing the overall customer experience
Diagnose and resolve data integration, ETL, and pipeline issues with guidance from senior team members, escalating complex issues as appropriate
Support the timely resolution of business-impacting data integration defects, failures, and processing issues to minimize disruptions to downstream reporting and analytics
Assist in root cause analysis activities and contribute to the implementation of corrective and preventive actions to improve platform reliability and operational efficiency
Create and maintain operational documentation, troubleshooting guides, and knowledge base articles to support ongoing platform operations and team knowledge sharing
Participate in a scheduled on-call rotation. Follow established troubleshooting and escalation procedures, working with senior engineers to resolve complex or business-critical incidents
Stay informed about emerging data engineering technologies, cloud platform capabilities, and new releases of tools used by the Analytics and Data Integration teams, while developing expertise in AWS, Databricks, and related technologies
Assist in evaluating new features, enhancements, and upgrades to data integration and analytics platforms, providing input on potential improvements and implementation opportunities
Support the administration, configuration, testing, and maintenance of data integration tools and cloud-based data platform solutions, including both custom-developed and vendor-provided software
Participate in the implementation, testing, and rollout of platform upgrades, patches, and new capabilities to ensure system stability and business continuity
Apply and promote established development, data engineering, security, and operational best practices to support reliable, scalable, and maintainable solutions
Collaborate with senior engineers and team members to identify process improvements, automation opportunities, and operational efficiencies that enhance platform performance and delivery quality
Continuously expand technical knowledge and skills through training, self-development, and hands-on experience with modern data engineering tools, frameworks, and cloud technologies
Occasional travel may be expected
Occasional work on nights and weekends to support production environment and meet project demands and deadlines
Qualification
Required
* Bachelor's degree in a quantitative or technical discipline such as Information Technology, Computer Science, Statistics, Economics, Mathematics, Engineering, or a related field
* 1-3 years of experience in software development, data engineering, ETL development, data integration, analytics, or a related technical discipline within large, multi-platform enterprise environments (Windows, Unix, Oracle)
* Exposure to or experience with data integration, ETL processes, data warehousing, or cloud-based data platforms
* Basic understanding of cloud platforms, preferably AWS, including services such as Amazon S3, IAM, and cloud-native data storage concepts
* Experience working with SQL and relational databases for data extraction, transformation, and analysis
* Basic programming experience in Python, SQL, or similar languages used for data engineering and analytics solutions
* Understanding of data warehousing concepts, data modeling fundamentals, and data integration best practices
* Familiarity with software development lifecycle (SDLC) processes and Agile development methodologies; experience with Jira or similar project management tools is a plus
* Strong analytical, problem-solving, and troubleshooting skills with the ability to learn new technologies quickly
* Ability to work effectively in a collaborative team environment and communicate technical concepts to both technical and non-technical stakeholders
* Demonstrated attention to detail and commitment to delivering high-quality, reliable solutions
* Working knowledge of Apache Spark (PySpark preferred) and distributed data processing concepts
* Experience developing data transformation and integration solutions using Python, SQL, or similar programming languages
* Familiarity with relational databases and SQL development; experience with Oracle databases and PL/SQL is a plus
* Understanding database concepts, data structures, data mapping, data transformations, and data modeling principles
* Exposure to Databricks and cloud-based data engineering technologies, including:
+ Delta Lake fundamentals
+ Databricks Workflows and Jobs
+ Data ingestion and transformation pipelines
* Experience or coursework involving batch data processing and data integration patterns, including:
+ Incremental data loading
+ Change Data Capture (CDC) concepts
+ Streaming data fundamentals
* Architecture and Platform Experience
+ Medallion Architecture (Bronze, Silver, Gold)
+ Batch and streaming data processing concepts
+ Data product and domain-oriented design concepts
+ Dimensional data modeling
+ Data quality and governance principles
+ Data lifecycle management
* Understanding of modern data platform concepts, including data lakes, data warehouses, and lakehouse architecture
* Familiarity with enterprise data architecture patterns such as:
* Basic understanding of:
* Contribute to scalable, maintainable, and reusable data solutions under the guidance of senior engineers
* Familiarity with software development best practices, including source control, code reviews, testing, and deployment processes
* Understanding of testing approaches for data pipelines and data quality validation
* Basic understanding of data security, access controls, and governance practices
* Familiarity with data lineage and compliance requirements in enterprise environments
* Proficient with Microsoft Office Suite (Word, Excel, PowerPoint)
* Familiarity with enterprise data platform / lakehouse platform and transactional system integrations
Preferred
**Although this is defined as a remote role and most work will be completed on a remote basis, the preferred candidate will live in the Chicago metropolitan area and can work from GATX's Chicago office on a periodic basis.**
* Familiarity with Databricks, Apache Spark, Delta Lake, or similar big data technologies is preferred
+ Exposure to Databricks Unity Catalog or similar data governance tools is preferred
* Familiarity with data visualization concepts and tools is preferred
* Relevant certifications (e.g., Databricks Certified Data Engineer, AWS Certified Data Analytics, Azure Data Engineer) are preferred
* Asset leasing industry experience is a plus, particularly within the Rail sector
* Working knowledge of Unix shell scripting is a plus
+ Understanding of cloud cost management, performance monitoring, and optimization concepts is a plus
* Familiarity with software development lifecycle (SDLC) processes and Agile development methodologies; experience with Jira or similar project management tools is a plus
* Familiarity with relational databases and SQL development; experience with Oracle databases and PL/SQL is a plus
* Occasional travel may be expected
* Occasional work on nights and weekends to support production environment and meet project demands and deadlines
Benefits
This role may be eligible to participate in the Company’s short-term incentive plan, the details of which will be provided to the applicant upon hire.
This is defined as a remote role and most work will be completed on a remote basis.
GATX is an equipment finance company based in Chicago, Illinois.