Data Engineer
- San José, Costa Rica
- Full-Time
- Remote
Job Description:
Job brief
We are seeking a Data Engineer to build, maintain, and improve data pipelines, transformation models, and infrastructure that support analytics, data science, and business operations.
In this role, you will work with a variety of internal and external data sources, develop reliable ingestion and transformation workflows, and help ensure that data is accurate, timely, scalable, and well documented.
You will collaborate closely with data scientists, analysts, engineers, and business stakeholders to deliver trusted data assets and improve the reliability and performance of the data platform.
Responsibilities
- Build and maintain data ingestion pipelines from internal and external sources, including APIs, SFTP, cloud storage, vendor feeds, and bulk files.
- Design, develop, and maintain data transformation models in a cloud data warehouse.
- Build and maintain tested and documented data models using dbt or similar transformation frameworks.
- Orchestrate and monitor scheduled data workflows using tools such as Prefect, Airflow, or Dagster.
- Implement monitoring and alerting to make pipeline failures visible and easier to diagnose.
- Investigate and resolve data quality, consistency, and reconciliation issues across multiple data sources.
- Optimize SQL queries, data models, and warehouse workloads for performance and cost efficiency.
- Partner with data scientists and analysts to turn one-off analyses into reliable production data assets.
- Review code and provide constructive feedback to improve code quality and engineering practices.
- Maintain clear documentation for data pipelines, transformations, dependencies, and data models.
- Communicate data sources, capabilities, and limitations to technical and nontechnical stakeholders.
- Design and implement proof-of-concept solutions for new tools, technologies, and data engineering approaches.
- Adopt AI-assisted development tools and methodologies where appropriate.
- Contribute to improvements in data architecture, platform reliability, scalability, and engineering processes.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Technology, or a related field, or equivalent practical experience.
- Experience building and maintaining data pipelines.
- Proficiency in SQL and Python.
- Experience working with cloud data warehouses such as Snowflake, BigQuery, or Redshift.
- Experience with data transformation frameworks such as dbt.
- Experience with workflow orchestration tools such as Prefect, Airflow, or Dagster.
- Good understanding of data modeling, ETL/ELT processes, and data transformation principles.
- Experience working with data from multiple sources and formats.
- Strong attention to detail when working with complex and potentially inconsistent datasets.
- Ability to identify edge cases, data quality issues, and potential technical risks.
- Familiarity with version control systems such as Git.
- Familiarity with containerization tools such as Docker.
- Ability to quickly learn and adopt new technologies.
- Strong analytical and problem-solving skills.
- Good communication and collaboration skills.
Nice to have
- Experience working with real estate, economic, geographic, public records, or similar datasets.
- Experience with entity resolution, record linkage, deduplication, or fuzzy matching.
- Experience with data warehouse cost management and query optimization.
- Experience with cloud platforms such as AWS or GCP.
- Experience with CI/CD pipelines and cloud infrastructure.
- Experience with BI and analytics platforms such as Sigma, Tableau, or Looker.
- Experience with AI-assisted coding tools.