Hi Pythonistas!
Your app has users.Users upload things.
- Profile pictures
- Post images
- Videos
- Documents
- Audio files
Where do these files go?
First instinct: save to your server's disk.
file.save('/var/www/uploads/profile_123.jpg')
Works locally.
Works for 100 users.
Then problems start.
Problem 1: Multiple Servers
Remember Post 6?
You have 3 app servers.
User uploads profile picture.
Saved on Server1's disk.
User visits their profile.
Load balancer routes to Server2.
Upload → Server1 (/var/uploads/pic.jpg) ✅
View → Server2 (/var/uploads/pic.jpg) ❌
File not there.
Broken.
Problem 2: Disk Space
App server disk: 500GB
1 million users × 5MB = 5TB of files
Doesn't fit.
Problem 3: Backups
Server crashes → files gone.
Complex to manage redundancy yourself.
Problem 4: Performance
Serving files uses CPU and memory.
Same resources handling your API.
Files competing with business logic.
Performance degrades.
The Solution: Object Storage
Don't store files on your server.
Store them in a dedicated system
built specifically for files.
Object Storage.
What Is Object Storage?
Stores data as objects.
Each object has:
Data → actual file content (bytes)
Key → unique identifier
Metadata → information about the file
Example:
Key: users/123/profile.jpg
Data: [binary image data]
Metadata: {
content-type: image/jpeg,
size: 245678,
uploaded-at: 2024-01-15
}
No folders. No hierarchy.
Just a flat key-value store for files.
Slashes in key simulate folders.
Underneath just a key and its value.
Amazon S3
Most popular object storage.
Simple Storage Service.
Launched 2006.
Changed how the world stores files.
Bucket:
Container for objects.
parseltongue-uploads ← bucket
parseltongue-backups ← another bucket
parseltongue-assets ← another bucket
Object:
A file inside a bucket.
parseltongue-uploads/users/123/profile.jpg
parseltongue-uploads/posts/456/image.png
Key:
Unique identifier within bucket.
users/123/profile.jpg
How S3 Works Under the Hood
Not a simple file system.
A massively distributed system.
When you upload:
- S3 receives the file
- Splits into chunks
- Stores across multiple servers
- Replicates across multiple data centers
- Returns URL to access
Amazon guarantees:
99.999999999% durability (11 nines)
11 nines means:
store 10 million files.
expect to lose 1 file every 10,000 years.
Practically impossible to lose data.
Using S3 in Python
import boto3
import os
# credentials from environment variables
# NEVER hardcode credentials
s3 = boto3.client(
's3',
aws_access_key_id=os.environ['AWS_ACCESS_KEY_ID'],
aws_secret_access_key=os.environ['AWS_SECRET_ACCESS_KEY'],
region_name='ap-south-1'
)
# upload
s3.upload_file(
'local_file.jpg',
'parseltongue-uploads',
'users/123/profile.jpg'
)
# download
s3.download_file(
'parseltongue-uploads',
'users/123/profile.jpg',
'downloaded.jpg'
)
Public vs Private Files
Public files:
Product images → anyone can view
Blog post images → anyone can view
s3.put_object_acl(
Bucket='parseltongue-uploads',
Key='users/123/profile.jpg',
ACL='public-read'
)
# direct URL, always accessible
url = 'https://parseltongue-uploads.s3.amazonaws.com/users/123/profile.jpg'
Private files:
User documents → only that user
Medical records → only authorized
Invoice PDFs → only that customer
Generate presigned URL for temporary access:
url = s3.generate_presigned_url(
'get_object',
Params={
'Bucket': 'parseltongue-uploads',
'Key': 'users/123/document.pdf'
},
ExpiresIn=3600 # valid 1 hour
)
URL expires after 1 hour.
After that → inaccessible.
This is how Google Drive share links work.
How Dropbox download links work.
The Big Security Question
Now you might think:
"For client-side uploads
do we put AWS credentials in the browser?"
Never. Ever.
const s3 = new AWS.S3({
accessKeyId: 'AKIAIOSFODNN7EXAMPLE', // ❌
secretAccessKey: 'wJalrXUtnFEMI/K7MDENG' // ❌
})
Anyone can:
open browser dev tools.
find your credentials.
delete all your files.
upload anything.
run up your AWS bill to millions.
This happens.
Real companies have been destroyed by this mistake.
The Safe Way: Presigned Upload URLs
Client never needs AWS credentials.
Ever.
AWS Credentials → live only on YOUR SERVER
→ never in browser
→ never in mobile app
How it works:
Client Your Server AWS S3
| | |
|-- "I want to upload" | |
| |-- credentials -->|
| |<- presigned URL -|
|<-- presigned URL ----| |
| | |
|------- uploads directly to S3 --------->|
| | |
|-- "upload done" ---->| |
| |-- save key ----->DB
Client gets a presigned URL.
Not your credentials.
Presigned Upload URL
def get_upload_url(request):
user_id = request.user.id # must be authenticated
url = s3.generate_presigned_url(
'put_object',
Params={
'Bucket': 'parseltongue-uploads',
'Key': f'users/{user_id}/profile.jpg',
'ContentType': 'image/jpeg',
'ContentLength': 500000 # max 500KB
},
ExpiresIn=300 # 5 minutes only
)
return {'upload_url': url}
Client uses it:
// get URL from YOUR server (no credentials)
const {upload_url} = await fetch('/api/upload/presigned-url', {
method: 'POST',
headers: {'Authorization': `Bearer ${userToken}`}
}).then(r => r.json())
// upload directly to S3
await fetch(upload_url, {
method: 'PUT',
body: file,
headers: {'Content-Type': 'image/jpeg'}
})
// tell server it's done
await fetch('/api/profile/picture', {
method: 'POST',
body: JSON.stringify({key: `users/${userId}/profile.jpg`})
})
Your server never sees the file.
S3 handles all bandwidth.
Fast. Scalable. Safe.
What If Someone Steals the Presigned URL?
Worst case:
they can upload one file.
to one specific key.
for 5 minutes.
That's it.
Cannot:
list your bucket ❌
delete files ❌
access other files ❌
run up your bill ❌
Much safer than credentials.
Should Files Go Through Server First?
Common question.
Short answer: No. They shouldn't.
Old way:
User → Server disk → S3
File touches your server.
Uses server RAM
Uses server disk
Uses server bandwidth
Large files block other requests
Modern way:
User → S3 directly (presigned URL)
User → Server (just metadata)
File never touches your server.
When server-side makes sense:
Virus scanning → must see file
Watermarking → must process file
Strict validation → must verify content
Even then process in memory:
import io
def upload_view(request):
file = request.FILES['photo']
# read into MEMORY (not disk)
file_bytes = file.read()
# validate in memory
if not is_valid_image(file_bytes):
return error_response
# upload bytes directly to S3
s3.upload_fileobj(
io.BytesIO(file_bytes),
'bucket',
'key'
)
No disk involved.
Memory only.
The rule:
Small file + no processing → presigned URL, direct to S3
Small file + processing → server memory → S3
Large file → presigned URL, direct to S3
Large file + processing → direct to S3 → async processing job
Server disk → almost never
Where to Store AWS Credentials
Never in code.
Never in git.
Environment variables (development):
export AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
export AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG
import os
key = os.environ['AWS_ACCESS_KEY_ID']
.env file:
AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG
Add .env to .gitignore.
Never commit to git.
IAM Roles (best practice on AWS):
Server runs on AWS?
Don't use credentials at all.
Assign IAM Role to your server.
Permissions handled automatically.
python
# no credentials anywhere
# IAM role handles it
s3 = boto3.client('s3')
AWS Secrets Manager (production):
import boto3, json
def get_credentials():
client = boto3.client('secretsmanager')
secret = client.get_secret_value(SecretId='myapp/s3')
return json.loads(secret['SecretString'])
The Mistake That Ruins Companies
This is real.
Happens regularly.
Developer pushes code to GitHub.
AWS credentials in the code.
Repository is public.
Bots scan GitHub every minute.
Find credentials within seconds.
Start mining Bitcoin on your AWS account.
Bill: $50,000 in one night.
Rule 1: Never put credentials in code
Rule 2: Never put credentials in git
Rule 3: Use environment variables
Rule 4: Use IAM roles on AWS (best)
Rule 5: Short expiry on presigned URLs
Rule 6: Add .env to .gitignore always
S3 + CDN
S3 alone serves from one region.
User in New York. S3 in Mumbai. Still slow.
Combine with CDN (Post 9):
S3 (origin) + CloudFront (CDN)
User requests image
↓
CloudFront nearest edge
↓
Cache hit → serve from edge (fast)
Cache miss → fetch from S3, cache at edge
S3 = storage.
CDN = delivery.
Together = fast global file serving.
Storage Classes
Different prices. Different access speeds.
Storage Class Use Case Cost
────────────────────────────────────────────────
S3 Standard Frequently accessed $$$
S3 Standard-IA Infrequent access $$
S3 Glacier Archives cents
S3 Glacier Deep Long-term backup very cheap
Lifecycle Rules: automate cost savings:
s3.put_bucket_lifecycle_configuration(
Bucket='parseltongue-uploads',
LifecycleConfiguration={
'Rules': [{
'Status': 'Enabled',
'Transitions': [
{'Days': 30, 'StorageClass': 'STANDARD_IA'},
{'Days': 90, 'StorageClass': 'GLACIER'}
],
'Expiration': {'Days': 365}
}]
}
)
Day 0: uploaded → Standard ($$$)
Day 30: moved → Standard-IA ($$)
Day 90: moved → Glacier (cents)
Day 365: deleted automatically
Old data costs less automatically.
Object Storage vs Others
File System → folders, hierarchy, your laptop disk
→ good for local dev, bad for scale
Block Storage → raw storage blocks (AWS EBS)
→ used by databases internally
→ not shareable across servers
Object Storage → flat key-value for files (S3)
→ good for images, videos, documents
→ bad for frequently updated files
→ never for databases
Other Providers
S3 is not the only option.
AWS S3 → most popular
Google Cloud Storage → tight GCP integration
Azure Blob Storage → Microsoft ecosystem
Cloudflare R2 → no egress fees
MinIO → self-hosted
DigitalOcean Spaces → simple, cheap
Most are S3 compatible.
Same API. Different endpoint.
# AWS S3
s3 = boto3.client('s3',
endpoint_url='https://s3.amazonaws.com')
# Cloudflare R2
s3 = boto3.client('s3',
endpoint_url='https://xxx.r2.cloudflarestorage.com')
# MinIO (self-hosted)
s3 = boto3.client('s3',
endpoint_url='http://localhost:9000')
Same code. Switch providers.
No code changes.
Complete Upload Flow
Here's the full picture:
- User selects file on frontend
- Frontend requests upload URL:
POST /api/upload/presigned-url
→ server validates authentication
→ server generates presigned URL
→ returns URL to frontend
- Frontend uploads directly to S3:
PUT [presigned URL]
[file bytes]
→ your server never sees the file
- Frontend tells server upload done:
POST /api/profile/picture
{key: 'users/123/profile.jpg'}
- Server saves key to database
- Async processing (if needed):
→ resize image
→ compress
→ upload processed versions to S3
- Frontend displays via CDN:
<img src="https://cdn.parseltongue.co.in/users/123/profile.jpg">
Mental Model
Object storage → flat key-value store for files
Bucket → container for objects
Object → file + metadata + key
S3 → most popular object storage
Presigned URL → temporary URL for private file access
Presigned upload → client uploads directly to S3
Public file → accessible via direct URL
Private file → only via presigned URL (expiring)
CDN + S3 → S3 stores, CDN delivers fast
Storage class → Standard, IA, Glacier (cost vs speed)
Lifecycle rules → auto-move between storage classes
IAM Role → best practice, no credentials in code
S3 compatible → same API, different providers
Never → credentials in client code
Never → files in database
Never → credentials in git
What Changed for Me
Before this:
I stored files on the server disk.
Moved to S3 only when disk was full.
After this:
files never touch my server.
client uploads directly to S3.
server only handles metadata.
CDN serves everything.
Infinitely scalable.
Zero file management headaches.
What's Coming Next
Phase 3: Storage is complete.
You now understand:
Post 13 → SQL vs NoSQL
Post 14 → Database Indexing
Post 15 → Object Storage
Next Phase 4: Reliability.
Your system handles scale.
But what happens when things break?
Servers crash.
Networks fail.
Databases go down.
How do you build a system
that survives all of this?
Message Queues.
The foundation of reliable systems.