Rafael Ignaulin
Senior Data Engineer
- rafa@ignaulin.com
- www.linkedin.com/in/rafa-ignaulin/
- Remote — Global (15+ countries, all major timezones)
Professional Summary
Senior Data Engineer with 5+ years building large-scale data platforms on Spark and Databricks — clickstream data foundations at billions of rows, CDC streaming, and production LLM systems (NL→SQL agents, MCP servers, RAG), working AI-native day to day with Claude and Cursor. Delivers for Nike and Inter&Co as an independent contractor through isTech, his own consultancy; previously full-time at PagSeguro/PagBank. Has delivered from 15+ countries across Europe, South America, and Oceania, always aligned to his clients' hours — location has never been a constraint on delivery.
Key Achievements
- Cut Spark data reads 98% at Nike (multi-terabyte per run down to tens of GB), cutting AWS costs 70%
- Automated Nike's A/B test measurement end-to-end with Airflow — weeks to minutes, ~50x more iterations
- Own the clickstream data foundation powering Nike's Marketing and Digital Experience analytics (billions of rows, Databricks/Spark/Airflow)
- Built an LLM analytics agent at Nike (Claude Sonnet, Streamlit, MLflow) that turns plain-English questions into SQL
- Own end-to-end a client segmentation and advisor-allocation platform at Inter&Co impacting 40M+ customers
- Founded isTech, a B2B data engineering consultancy delivering to US enterprise clients across 6+ timezones
Experience
Senior Data Engineer (Data Foundation, Marketing & Digital Experience) · Nike (Contract via isTech)
- Cut Spark data reads 98% (multi-terabyte per run down to tens of GB), cutting AWS costs 70%.
- Automated A/B test measurement with Airflow: 99.5% faster, weeks to minutes, ~50x more iterations.
- Own the clickstream fact table (billions of rows, hundreds of GB, millions of rows added daily) powering Marketing and Digital Experience analytics on Databricks/Spark/Airflow.
- Built an LLM analytics agent (Claude Sonnet, Streamlit, MLflow) turning plain-English questions into SQL, with confidence scoring on every answer.
- Built a FastMCP/FastAPI server exposing a dozen-plus analytics tools to Cursor via YAML-based tool registration — no code changes to add a tool.
- Contributed a dozen loader classes to the team's Databricks/Airflow job library (task orchestration, Delta load/validate, A/B testing, freshness checks).
- Debugged a large production analytics API: root-caused routing/table/RLS bugs and built a cache-validation test suite covering dozens of metrics.
Data Engineering Lead (Client Segmentation & Advisor Allocation) · Inter&Co (Contract via isTech)
- Own end-to-end a client segmentation and advisor-allocation platform (Spark, Kafka, Delta Lake) impacting 40M+ customers at a digital bank.
- Built a 6-zone medallion lakehouse on S3/Delta Lake, orchestrated by Airflow across EMR and Kubernetes, across 8 repositories.
- Designed CDC ingestion (Kafka/MSK, Debezium) and a real-time update engine keeping segmentation current every 15-20 minutes.
- Built a reusable data-quality framework adopted across every ingestion pipeline on the platform.
- Designed the advisor-allocation algorithm and set the engineering standards (testing, CI/CD, alerting) across the team's 8 repositories.
- Progressed Mid-Level → Senior → Lead across the engagement.
Founder & Principal Data Engineer · isTech (Own Consultancy)
- Founded the B2B consultancy all contract work above runs through — Data Warehouses, Lakes, and Lakehouses on AWS/Azure with Spark, Kafka, Airflow, Terraform.
- Delivered from 15+ countries across 6+ timezones while holding enterprise contracts, always aligned to client hours.
Data Engineer (Junior → Mid-Level) · PagSeguro / PagBank (Full-time)
- Designed and tuned Amazon Redshift Data Warehouse with star/snowflake schemas, optimizing queries across 20+ tables with billions of rows for 10x performance gain.
- Promoted Junior → Mid-Level; transitioned out in Oct 2024 as isTech contract work became the focus.
Skills
Languages
Python, SQL, Scala
Big Data & Streaming
Spark, PySpark, Kafka, Flink, Debezium, Confluent Schema Registry
Cloud
AWS (S3, EMR, Redshift, DynamoDB, Lambda), Azure, GCP
Data Platforms
Databricks, Snowflake, Delta Lake, Unity Catalog
Orchestration & IaC
Airflow, Dagster, Terraform, Docker, Kubernetes
Databases
Redshift, PostgreSQL, DynamoDB, Trino, Presto, Athena
AI & Agentic
MCP Servers, FastMCP, LLM Integration, RAG, Prompt Engineering, Cursor AI, LangChain, MLflow
BI & Visualization
Tableau, Databricks AI/BI, Power BI, Streamlit
APIs & Frameworks
FastAPI, asyncpg, httpx
Education
M.S. Computer Science (OMSCS) · Georgia Institute of Technology
One of the world's top-ranked online CS master's programs
Bachelor in Computer Science · Unochapecó
Student Honor Medal — Top 1% of graduating class
Certifications
- AWS Certified Solutions Architect – Professional
- AWS Data Engineer Associate
- AWS Black Belt
- Microsoft Azure Solutions Architect Expert (AZ-305)
- Databricks Data Engineer Professional
13 professional certifications in total across AWS, Databricks, Azure, Airflow, and Oracle.
Languages & Work Authorization
- English — C2 — Professional Proficiency
- Portuguese — Native
Brazilian citizen — eligible for US and EMEA work visas