To use an S3-compatible store from code you need four things beyond the access key and secret: the endpoint URL, a region name, path-style addressing switched on, and - with SDKs released since early 2025 - checksums set to "when required". In Python that is boto3.client("s3", endpoint_url=..., region_name="us-east-1", config=Config(s3={"addressing_style": "path"})). In Node it is new S3Client({ endpoint, region: "us-east-1", forcePathStyle: true }) from @aws-sdk/client-s3. Everything after that - uploads, downloads, listing, deleting - is the same code you would write for Amazon S3. This post sets both SDKs up properly, then works through the operations an application actually needs, the large-file cases, and the errors you will meet.
The settings every client needs#
| Setting | Python (boto3) | Node (@aws-sdk/client-s3) | Why |
|---|---|---|---|
| Endpoint | endpoint_url= | endpoint: | Otherwise the SDK talks to AWS |
| Region | region_name= | region: | Part of the request signature |
| Path-style | Config(s3={'addressing_style': 'path'}) | forcePathStyle: true | Bucket in the path, not the hostname |
| Credentials | aws_access_key_id, aws_secret_access_key | credentials: {...} | Or environment variables |
| Checksums | request_checksum_calculation="when_required" | requestChecksumCalculation: "WHEN_REQUIRED" | Compatibility with non-AWS servers |
Path-style is the setting people forget. Without it, the SDK sends the request to bucket-name.your-endpoint, which has no DNS record, and the error says the host could not be found. Path-style vs virtual-hosted addressing explains why S3-compatible endpoints use path-style and what the region name really does.
The checksum settings are newer. In January 2025 the AWS SDKs started adding CRC integrity checksums to uploads by default (boto3 from 1.36, the JavaScript SDK from 3.729). Many S3-compatible servers do not recognise the new headers, and the result is failed uploads with SignatureDoesNotMatch, MissingContentLength or XAmzContentSHA256Mismatch while reads keep working. Setting both options to "when required" restores the older behaviour. If your store accepts the default, you can leave them out; setting them costs nothing either way.
Python: a boto3 client that works#
Install with pip install boto3. Keep configuration in environment variables rather than in the code - environment variables and secrets covers why and how.
import osimport boto3from botocore.config import Configs3 = boto3.client( "s3", endpoint_url=os.environ["S3_ENDPOINT"], # https://s3.example.com region_name=os.environ.get("S3_REGION", "us-east-1"), aws_access_key_id=os.environ["S3_ACCESS_KEY_ID"], aws_secret_access_key=os.environ["S3_SECRET_ACCESS_KEY"], config=Config( signature_version="s3v4", s3={"addressing_style": "path"}, request_checksum_calculation="when_required", response_checksum_validation="when_required", retries={"max_attempts": 5, "mode": "standard"}, ),)BUCKET = os.environ["S3_BUCKET"]A few notes on the choices:
signature_version="s3v4"is the default in current botocore; stating it guards against old installs that defaulted to SigV2 for some operations.- The
retriesblock uses botocore'sstandardmode, which retries throttling and transient network errors with backoff. The defaultlegacymode retries fewer cases. - A boto3 client is thread-safe. Create one per process and reuse it. Creating a client per request costs tens of milliseconds and a new connection pool each time.
- The checksum options in
Configneed botocore 1.36 or newer. On an older version, remove those two lines; the old version does not send the new checksums anyway.
If you prefer, the same values can come from the standard variables AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_DEFAULT_REGION and AWS_ENDPOINT_URL_S3, in which case boto3.client("s3", config=...) picks them up with no arguments. Explicit arguments are easier to read when the same process also talks to real AWS for something else.
Python: uploads, downloads and listing#
The operations an application uses day to day:
# Upload a local file. Handles multipart automatically above 8 MB.s3.upload_file("report.pdf", BUCKET, "reports/2026/report.pdf", ExtraArgs={"ContentType": "application/pdf"})# Upload bytes you already have in memorys3.put_object(Bucket=BUCKET, Key="notes/hello.txt", Body=b"hello", ContentType="text/plain; charset=utf-8")# Download to a file, or read into memorys3.download_file(BUCKET, "reports/2026/report.pdf", "/tmp/report.pdf")body = s3.get_object(Bucket=BUCKET, Key="notes/hello.txt")["Body"].read()# Does it exist?from botocore.exceptions import ClientErrortry: s3.head_object(Bucket=BUCKET, Key="notes/hello.txt")except ClientError as e: if e.response["Error"]["Code"] in ("404", "NoSuchKey", "NotFound"): print("missing") else: raise# Deletes3.delete_object(Bucket=BUCKET, Key="notes/hello.txt")Set ContentType on upload. S3 stores whatever you give it and returns it on download; if you leave it out, objects come back as binary/octet-stream and browsers download images instead of showing them.
Listing is where people write bugs. list_objects_v2 returns at most 1,000 keys per call, and code that calls it once silently ignores everything after the first thousand. Use the paginator:
paginator = s3.get_paginator("list_objects_v2")total = 0for page in paginator.paginate(Bucket=BUCKET, Prefix="reports/2026/"): for obj in page.get("Contents", []): total += obj["Size"] print(obj["Key"], obj["Size"], obj["LastModified"])print(f"{total / 1e6:.1f} MB")Pass Delimiter="/" to list one "folder" level at a time: keys below the next slash are folded into CommonPrefixes. S3 has no real folders - a key is one flat string - but prefixes and delimiters let you browse as if it did.
upload_file and download_file use the transfer manager, which splits large files into parts and moves them in parallel. Its defaults are an 8 MB threshold, 8 MB parts and 10 threads. On a small server with little memory, lower the concurrency rather than raising it:
from boto3.s3.transfer import TransferConfigcfg = TransferConfig(multipart_threshold=16 * 1024 * 1024, multipart_chunksize=16 * 1024 * 1024, max_concurrency=4)s3.upload_file("world-backup.tar.zst", BUCKET, "backups/world.tar.zst", Config=cfg)Each in-flight part is held in memory, so 4 threads of 16 MB is roughly 64 MB of buffers. That matters on a 1 GB app plan.
boto3 inside async frameworks
boto3 is synchronous. Called directly from an async def view in FastAPI, Starlette or an async Django view, every upload blocks the event loop and every other request on that worker waits for it. Either declare the endpoint with plain def (FastAPI then runs it in a thread pool for you), or push the call into a thread explicitly:
import asyncioasync def save_avatar(data: bytes, key: str) -> None: await asyncio.to_thread( s3.put_object, Bucket=BUCKET, Key=key, Body=data, ContentType="image/png" )Third-party async wrappers exist (aiobotocore and aioboto3), but they track botocore versions closely and lag behind it. For an application doing a handful of uploads per second, a thread is simpler and perfectly adequate.
Node: an SDK v3 client that works#
The JavaScript SDK v3 is modular: install only what you use. For most applications that is two or three packages.
$ npm install @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-request-presignerimport { S3Client } from "@aws-sdk/client-s3";export const s3 = new S3Client({ endpoint: process.env.S3_ENDPOINT, // https://s3.example.com region: process.env.S3_REGION ?? "us-east-1", forcePathStyle: true, credentials: { accessKeyId: process.env.S3_ACCESS_KEY_ID, secretAccessKey: process.env.S3_SECRET_ACCESS_KEY, }, requestChecksumCalculation: "WHEN_REQUIRED", responseChecksumValidation: "WHEN_REQUIRED", maxAttempts: 5,});export const BUCKET = process.env.S3_BUCKET;As with boto3, create the client once at module level and import it everywhere. SDK v2 (aws-sdk) reached end of support in September 2025; if you are maintaining code that still uses it, the equivalent options were s3ForcePathStyle: true and endpoint, and migrating is worth doing for the smaller install and the maintained code alone.
Node: uploads, downloads and listing#
Every operation is a command object passed to s3.send():
import { PutObjectCommand, GetObjectCommand, HeadObjectCommand, DeleteObjectCommand, paginateListObjectsV2 } from "@aws-sdk/client-s3";import { readFile } from "node:fs/promises";// Upload a small file from memoryawait s3.send(new PutObjectCommand({ Bucket: BUCKET, Key: "notes/hello.txt", Body: "hello", ContentType: "text/plain; charset=utf-8",}));// Download: Body is a stream with helper methodsconst res = await s3.send(new GetObjectCommand({ Bucket: BUCKET, Key: "notes/hello.txt" }));const text = await res.Body.transformToString();// Exists?try { await s3.send(new HeadObjectCommand({ Bucket: BUCKET, Key: "notes/hello.txt" }));} catch (err) { if (err.name !== "NotFound" && err.$metadata?.httpStatusCode !== 404) throw err;}// List everything under a prefix, all pagesfor await (const page of paginateListObjectsV2({ client: s3 }, { Bucket: BUCKET, Prefix: "reports/" })) { for (const obj of page.Contents ?? []) console.log(obj.Key, obj.Size);}await s3.send(new DeleteObjectCommand({ Bucket: BUCKET, Key: "notes/hello.txt" }));The response Body in Node is a readable stream extended with transformToString() and transformToByteArray(). For a large object, do not buffer it - pipe it:
import { createWriteStream } from "node:fs";import { pipeline } from "node:stream/promises";const { Body } = await s3.send(new GetObjectCommand({ Bucket: BUCKET, Key: "backups/world.tar.zst" }));await pipeline(Body, createWriteStream("/tmp/world.tar.zst"));One trap: if you request an object and never read or destroy the body, the socket stays checked out of the connection pool. Do that fifty times (the default maxSockets of the SDK's HTTP agent is 50) and every later request hangs waiting for a free socket. Always consume or destroy Body.
Large uploads and streams in Node#
PutObjectCommand with a stream of unknown length fails, because a single PUT needs a Content-Length. For files of unknown size, or anything large, use Upload from @aws-sdk/lib-storage, which splits the input into multipart parts:
import { Upload } from "@aws-sdk/lib-storage";import { createReadStream } from "node:fs";const upload = new Upload({ client: s3, params: { Bucket: BUCKET, Key: "backups/world.tar.zst", Body: createReadStream("world.tar.zst") }, queueSize: 4, // parts in flight partSize: 16 * 1024 * 1024, // 16 MB, minimum 5 MB leavePartsOnError: false,});upload.on("httpUploadProgress", (p) => console.log(p.loaded, p.total));await upload.done();Multipart uploads have rules that come from the S3 API itself: every part except the last must be at least 5 MB, and an upload may have at most 10,000 parts. With 16 MB parts that caps a single object at about 160 GB; raise partSize for anything bigger. A multipart upload that is started and never completed or aborted leaves its parts on the server, using space. leavePartsOnError: false aborts on failure, but a process that is killed mid-upload cannot clean up after itself. List stray uploads with ListMultipartUploadsCommand (or aws s3api list-multipart-uploads) once in a while and abort the old ones - on storage without lifecycle rules, nothing else will.
Designing keys, metadata and headers#
The SDK will store anything under any key, which is exactly why the key layout deserves five minutes of thought before the first upload. A few habits save trouble later.
Prefix by what you will list or delete together. If you will one day need "everything for user 1042" or "every export older than March", put that in the front of the key: users/1042/avatar.png, exports/2026/03/14/orders.csv. Listing by prefix is cheap; finding scattered keys means listing the whole bucket and filtering in code.
Never use the user's file name as the key. Two users upload photo.jpg, and the second overwrites the first. File names also bring spaces, Unicode, slashes and .. with them. Generate the key yourself - a UUID or a content hash plus the right extension - and keep the original name in your database or in object metadata if you need to show it.
Keep keys URL-friendly. Any UTF-8 string up to 1,024 bytes is legal, but keys end up in URLs, logs and shell commands. Lowercase letters, digits, hyphens, underscores, dots and slashes cause no trouble anywhere.
Object metadata travels with the object. Standard HTTP headers - ContentType, CacheControl, ContentDisposition - are returned to whoever downloads the object, and custom metadata goes in Metadata as string pairs, returned as x-amz-meta-* headers:
s3.upload_file("avatar.png", BUCKET, "users/1042/avatar-3f9c.png", ExtraArgs={ "ContentType": "image/png", "CacheControl": "public, max-age=31536000, immutable", "Metadata": {"original-name": "my photo.png", "uploaded-by": "1042"},})ContentDisposition: attachment; filename="report.pdf" makes browsers download rather than display, which is what you want for user-supplied files you do not trust to render. A long CacheControl is safe when the key changes whenever the content does - which is another reason to put a hash or version in the key. HTTP caching headers explained covers the values in detail.
Metadata cannot be edited in place. Changing it means copying the object onto itself with new metadata (copy_object with MetadataDirective="REPLACE"), so set it right at upload time.
Errors you will meet#
| Error | Meaning | Fix |
|---|---|---|
ENOTFOUND bucket.s3.example.com | Virtual-hosted style | forcePathStyle: true / addressing_style: path |
SignatureDoesNotMatch on every call | Wrong secret, or a proxy changing Host | Re-copy the secret; check the proxy |
SignatureDoesNotMatch on uploads only | New default checksums | Checksum options "when required" |
InvalidAccessKeyId | Key not known to this endpoint | Check you are not hitting AWS by accident |
NoSuchBucket | Bucket name typo, or style mismatch | Check the name; check path-style |
RequestTimeTooSkewed | Clock off by more than 15 minutes | Fix NTP on the client machine |
EPROTO / wrong version number | HTTPS spoken to a plain-HTTP port | Use http:// for the raw port or the HTTPS hostname |
InvalidAccessKeyId deserves a second look when it appears. The commonest cause is not a wrong key but a missing endpoint: an environment variable was not loaded, the SDK fell back to Amazon, and Amazon has never heard of your key. Log the client's resolved endpoint at startup and the mystery disappears.
When an error is not obvious, turn on the SDK's own logging for one run. In boto3, boto3.set_stream_logger("botocore", logging.DEBUG) prints every request with its URL and headers, which shows at a glance whether the endpoint, style and region are what you expected. In Node, pass logger: console to the S3Client constructor for a lighter version, or inspect err.$metadata and err.$response on the caught error, which carry the HTTP status and the raw response from the server. Remove both before production; debug output includes headers you do not want in logs.
RequestTimeTooSkewed is the other surprise. Signed requests carry a timestamp and servers reject any more than 15 minutes from their own clock. It almost never happens on a hosted server, and often happens on a laptop that has been asleep or a container with no time sync.
Using it from an app server#
The usual reason an app talks to S3 is to keep user files off the app's own disk: uploads, generated reports, exports. That has two benefits on a small server - the app's disk stays small and fast, and redeploying or rebuilding the app does not touch the files.
On RE:NODE the pieces fit like this. The Node.js and Python app lines run your code; the S3 storage line gives you an endpoint with an access key and secret generated for the server and a first bucket already created. Put the endpoint, keys and bucket in the app's environment variables on the Startup tab, use the HTTPS hostname from the storage plan's proxy slot as S3_ENDPOINT, and the code above runs unchanged. Plain HTTP on the storage port works for a quick test, but objects and request bodies cross the network unencrypted, so production traffic should go through the HTTPS name.
Use a separate bucket for each environment - media-dev and media-prod, say - selected by the S3_BUCKET variable. A test suite that cleans up after itself by deleting a prefix should never be one typo away from the production bucket. Remember also that the storage line keeps one copy of each object on NVMe in one location, without replication: it is a sensible home for user files, but anything you cannot recreate should also be copied somewhere else.
When users need to upload directly from the browser, do not stream the file through your app at all: hand the browser a presigned URL and let it talk to storage. Presigned URLs for uploads covers that pattern in both languages. And if what you are storing is backups rather than user files, a dedicated tool is a better fit than hand-written code - restic backups to S3 handles encryption, deduplication and retention for you.
FAQ#
Do I need the full aws-sdk package in Node?
No. Version 3 is split per service; @aws-sdk/client-s3 is all you need for basic operations, plus @aws-sdk/lib-storage for multipart uploads and @aws-sdk/s3-request-presigner for presigned URLs. The old monolithic aws-sdk v2 is out of support.
Can I use the same code against AWS and a self-hosted store?
Yes, if the endpoint, region and path-style flag come from configuration. Leave the endpoint empty and path-style off for AWS; set them for the custom endpoint. The API calls themselves are identical.
Why do uploads fail but downloads work?
Most often the default checksums added in 2025 SDK releases. Set the request checksum calculation and response checksum validation options to "when required". If that does not fix it, check the object size against the 5 GB limit for a single PUT and switch to multipart.
How do I create a bucket from code?
s3.create_bucket(Bucket="name") in boto3 or CreateBucketCommand in Node. On AWS outside us-east-1 you must also pass a location constraint; on an S3-compatible store it is usually ignored. Many applications are simpler if the bucket is created once by hand and the code only uses it.
Is it safe to put the secret key in my code?
No. Put it in environment variables or a secrets file outside the repository. A secret committed to Git should be treated as published, even in a private repository, and replaced.




Comments
Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.