We work on enterprise data platforms and scalable pipelines on Google Cloud. We build reliable systems for ingesting, transforming, and analyzing large volumes of data.
Data Engineer - Big Data & GCP
Data Engineer / Big Data
Your role:
- Design and develop batch and streaming data processing pipelines using Apache Beam and Apache Spark on Google Cloud Dataproc.
- Model, optimize, and query large datasets in BigQuery, with a focus on cost and performance.
- Work with BigQuery Studio to explore data, develop analytical notebooks, and collaborate with data science and analytics teams.
- Integrate heterogeneous data sources (relational databases, APIs, event streams) into GCP pipelines.
- Monitor data quality and implement tests and alerts on production pipelines.
- Work with the Data Science team to ensure that data is accessible, reliable, and well-documented.
- Help define the team's data engineering standards: naming conventions, data catalog, and data lineage.
Your profile:
- Apache Beam/Spark — Pipeline Development — Intermediate Level
- Google Cloud Dataproc — Clusters, jobs, tuning — intermediate level
- BigQuery — Advanced SQL, Optimization — Intermediate/Advanced Level
- BigQuery Studio — Notebooks, exploration — beginner/intermediate level
- Python — Pipelines and Scripting — Intermediate Level
- Advanced SQL — Window Functions, CTE, Optimization — Intermediate Level
- Git/CI-CD — Collaborative Workflow — Intermediate Level
Nice to have:
- Experience with Google Cloud Dataflow (fully managed Beam pipelines).
- Knowledge of pipeline orchestrators: Apache Airflow / Cloud Composer.
- Introduction to data modeling: star schema, snowflake schema, Data Vault.
- Google Cloud Professional Data Engineer certification (or currently pursuing it).
What we offer:
- Approx. €30,000
- Permanent contract (National Collective Bargaining Agreement for the Retail Sector).
- Company canteen
- Structured professional development plan.

