I'm a Data Engineer focused on building scalable, reliable, and production-ready data platforms. I work across cloud data engineering, distributed processing, batch and streaming pipelines, data modeling, and modern data lakehouse architectures.
I enjoy solving problems where data volume, reliability, performance, and scalability actually matter.
- Data Engineer with experience building and maintaining cloud-based data pipelines.
- Strong focus on AWS, Azure, Apache Spark, PySpark, SQL, and distributed data processing.
- Interested in streaming systems, data platforms, lakehouse architecture, and large-scale data engineering.
- Currently deepening my expertise in Databricks, Spark, Kafka, Delta Lake, data modeling, and system design.
- Preparing for opportunities where I can work on challenging, large-scale data systems.
- Python
- SQL
- Java
- PySpark
- Apache Spark
- Spark Structured Streaming
- Apache Kafka
- ETL / ELT
- Batch Processing
- Data Modeling
- Data Warehousing
- Data Lakehouse Architecture
AWS
- S3
- Glue
- EMR
- Athena
- Lambda
- Step Functions
- Redshift
- RDS
- ECR
Azure
- ADLS Gen2
- Azure Data Factory
- Azure Databricks
- Azure Synapse
- Unity Catalog
- Azure Purview
- PostgreSQL
- Oracle
- MongoDB
- Redshift
- Delta Lake
- Parquet
- Git
- GitHub Actions
- Docker
- Kubernetes
- Jenkins
- ArgoCD
- CloudWatch
- Splunk
- Advanced Apache Spark internals and optimization
- Spark Structured Streaming
- Kafka and event-driven architectures
- Delta Lake and Lakehouse architecture
- Databricks
- Advanced SQL and data modeling
- Data-intensive system design
- Distributed systems
- AWS and Azure data platforms
Building scalable batch and streaming pipelines using cloud-native and distributed technologies.
Working with Spark and PySpark to process large datasets efficiently while understanding partitioning, shuffles, joins, caching, and performance optimization.
Exploring real-time data pipelines using Kafka and Spark Structured Streaming.
Practicing dimensional modeling, event modeling, state modeling, interval modeling, temporal modeling, and other analytical data patterns.
Designing and working with modern data architectures across AWS and Azure.
- LinkedIn: Akshay V Doyizode
- GitHub: akshayvdoizode
I'm interested in data engineering, distributed systems, cloud platforms, and the engineering challenges that come with building reliable data products at scale.

