
Stop Paying Too Much for Spark S3 Reads
If you’re running Spark jobs on AWS and your S3 read costs keep climbing, you’re probably hitting some classic performance traps that are easy to miss but expensive to ignore.
This guide is for data engineers and ML engineers who work with Apache Spark on AWS — whether that’s on EMR, Glue, or a self-managed cluster — and want to cut down on slow, costly S3 reads without rewriting everything from scratch.
We’ll walk through how Spark actually reads data from S3 (and why that matters more than you’d think), which expensive S3 read patterns are quietly killing your job performance, and how to fix your data layout and Spark AWS configuration to get faster, cheaper reads. We’ll also look at some AWS-native tools that do a lot of the heavy lifting for you.
By the end, you’ll have a clear picture of where the waste is coming from and what to do about it.
Understanding How Spark Reads Data from S3

How S3 Differs from a Traditional Filesystem
S3 is object storage, not a filesystem. Every “folder” is just a key prefix.
Why S3 Read Latency Hurts Spark Job Performance
High API call volume kills performance. Each small file triggers a separate GET request.
Common Misconceptions That Lead to Costly Mistakes
- Treating S3 like HDFS
- Ignoring file count impact on Spark S3 optimization
Identifying Expensive S3 Read Patterns to Avoid

A. Small File Problem
Thousands of tiny files force Spark to make separate S3 API calls per file, killing performance.
B. Repeated Full Scans
Reading entire datasets without filters wastes compute and spikes S3 costs fast.
C. Unpartitioned Data
Without partitioning, every query scans everything unnecessarily.
D. Shuffle-Heavy Operations
Joins and aggregations repeatedly pull data from S3, creating expensive read amplification.
Optimizing Data Layout on S3 for Faster Reads

A. Partition Your Data Strategically
Partition by high-cardinality filter columns like date or region to skip irrelevant files entirely.
B. Use Parquet
Parquet’s column pruning slashes S3 read volume dramatically.
C. Right-Size Files
Target 128–256MB files to balance parallelism without excessive overhead.
D. Bucketing
Eliminates shuffle reads on joins.
E. Compress Smartly
Use Snappy for speed.
Configuring Spark and AWS for Maximum S3 Read Efficiency

A. Enable S3A Committer to Speed Up Write and Read Cycles
Switch from the default FileOutputCommitter to the S3A Magic Committer to cut rename overhead drastically.
B. Tune Spark Parallelism Settings to Match S3 Throughput
Set spark.sql.shuffle.partitions and maximizeResourceAllocation to align with your S3 request limits.
C. Use AWS Glue Data Catalog for Efficient Partition Pruning
Glue pushes partition filters early, slashing unnecessary S3 scans automatically.
Leveraging AWS-Native Tools to Reduce S3 Read Costs

A. Use S3 Intelligent-Tiering
Automatically moves frequently read data to lower-latency storage tiers, cutting costs without manual intervention.
B. Cache Hot Datasets
Amazon FSx for Lustre dramatically speeds up repeated Spark S3 reads by caching hot datasets locally.
C. Apply S3 Select
Push filters directly to S3—Spark retrieves only needed rows, slashing data transfer and read costs.
D. Monitor with Cost Explorer
Track expensive S3 read patterns using S3 Storage Lens alongside AWS Cost Explorer for actionable spending insights.

Getting Spark to play nicely with S3 isn’t rocket science, but it does require some intentional decisions. From understanding how Spark pulls data off S3 to spotting the read patterns that quietly drain your budget, small tweaks in how you structure your data and configure your setup can make a massive difference in both speed and cost.
Start by cleaning up your data layout, tuning your Spark and AWS settings, and leaning on AWS-native tools that are built to make this easier. If you put even a handful of these strategies into practice, you’ll likely notice faster job times and a much friendlier AWS bill. Pick one area to tackle first, get comfortable with it, and build from there.














