English

Developer tools · Base64 encoder & decoder

Base64 in JSON APIs: why binary fields get encoded and what it costs you

· Why it matters

base64 encoding

A JSON field containing a long Base64 string representing binary data encoded for transport
Original ToolAcre vector illustration

JSON has no byte type, so binary data is usually Base64-encoded into a string. This post explains why that convention exists, what it costs in size and CPU, and when a separate binary endpoint is the better call.

The PDF field that dominated the response — a concrete API payload where one Base64 blob outweighed everything else

An API response contains a large object with a single field that dominates the payload size. The response is JSON, so every value is a string or number. Most fields are small: user IDs, timestamps, status codes. One field contains imageData or fileContents and is a 40-kilobyte Base64 string. The entire response is 50 kilobytes. That single field accounts for 80 percent of the transfer, which seems wasteful because the server sent it as bytes originally and the client needs bytes again eventually.

Base64 solved this problem: JSON has no native byte type, so binary data must be wrapped in a string. Base64 converts arbitrary bytes into ASCII characters that are safe in JSON. Both the client and server must encode on send and decode on receive, adding CPU overhead. The resulting payload is about a third larger than the raw bytes. This post explains why that convention exists, what it costs in practice, and when breaking the JSON constraint with a separate binary endpoint is worth it.

Why JSON cannot carry raw bytes — strings must be valid Unicode text, so arbitrary bytes need a textual wrapper

JSON is a text format where all values must be valid Unicode text. The specification defines strings, numbers, booleans and null. It does not have a byte array or buffer type. If an API needs to return binary data like an image, a cryptographic signature or a file upload, it cannot put the raw bytes into the JSON object directly. The bytes might contain characters that JSON parsers interpret as structural markers. A null byte in the middle of a binary blob could terminate a string early or break the parser.

The common solution is to encode the binary data as Base64, producing a string of ASCII characters that JSON parsers treat as plain text. The receiving client then decodes the Base64 back to bytes and uses them. This encoding step happens at the API level, hidden from most developers, but it is a real cost that accumulates when APIs return many binary fields. The costs of Base64 in JSON compound across the request-response cycle. The size penalty is the first: Base64 output is approximately 33 percent larger than its input due to the encoding overhead.

The costs: a third more bytes, decode time and memory copies — where each cost appears in a typical client

A 30-megabyte video file becomes 40 megabytes when Base64-encoded. Downloading 40 instead of 30 megabytes costs bandwidth and battery on mobile devices, and time for users on slow connections. The second cost is CPU time. The server must encode the binary data as Base64 before it can be stringified into JSON. The client must parse the JSON and then decode each Base64 field back to bytes. For a response with multiple binary fields or a client that processes thousands of responses, that CPU time accumulates.

On constrained devices like phones, JavaScript string operations and the TextDecoder used for Base64 decoding consume battery and slow down the application. The third cost is memory: the JSON parser creates a string object for the Base64 field, then decoding creates another copy as a Uint8Array. A large field is instantiated twice in memory before the application can use it. A worked example clarifies the cost. Suppose an API endpoint returns user profile data including a 100-kilobyte avatar image. The server reads the image from disk as bytes, encodes it to Base64, and includes it in the JSON response.

Worked example: inspecting a Base64 field from an API response — decoding it in the browser to confirm what the server actually sent

The JSON response is now about 135 kilobytes (33% overhead plus the other fields). The client downloads 135 kilobytes instead of 100. In the browser, the JSON parser creates a JavaScript string object for the Base64 data.

When the application needs the image, it calls the Base64 decoder, which creates a Uint8Array of the original 100 kilobytes. For the few milliseconds of decoding, both objects exist in memory. If the page displays ten profiles with avatars, the cost is multiplied. The alternative is for the API to return the JSON response with a separate URL for each avatar resource, letting the browser handle the image downloads with its native caching, progressive rendering and memory management.

Alternatives: multipart, separate download URLs and raw binary endpoints — the trade-offs of each

The trade-off between including binary in JSON and fetching it separately depends on the API purpose and usage patterns. For a search results page that returns hundreds of tiny thumbnail images, fetching each as a separate request defeats HTTP connection pooling and caching. Inlining them as Base64 in the JSON response might be faster. For a detailed profile page that requests one or two high-resolution images, separate downloads are clearly better. API documentation should state the maximum size of Base64 fields and when clients should expect separate endpoints.

If a field regularly exceeds a kilobyte or two, the inline-Base64 strategy is a sign that the API design needs reconsideration. Alternatives to Base64 in JSON exist but each has trade-offs. Multipart MIME responses separate binary and text so the binary section is sent as raw bytes and only the text section is JSON. This requires the client to parse a multipart message instead of just calling JSON.parse, adding complexity. A separate download URL in the JSON response points the client to fetch the binary resource separately.

Conventions worth stating in your API docs — standard versus base64url alphabet, padding and maximum sizes

This works well when the binary resource is large or accessed less frequently than the metadata. A raw binary endpoint that returns only bytes and abandons JSON entirely is the simplest approach but removes the structure that JSON provides. Some APIs return compressed binary data and Base64-encode that, reducing the size penalty but adding decompression overhead. The choice depends on the expected usage: small fields are fine inline, large fields belong in separate resources, and structured data is worth keeping in JSON even with the Base64 cost.

Conventions matter for interoperability. APIs that Base64-encode binary data should document it clearly and state whether the alphabet is standard or URL-safe. Standard Base64 uses + and /, which are safe in JSON strings but not in URLs. URL-safe Base64 replaces them with - and _, which is appropriate for data: URIs but unnecessarily escaped in JSON. The documentation should specify whether padding is included or omitted, because both are valid Base64 but a client that expects padding and receives unpadded data will fail silently or produce garbage.

What this does not cover — protobuf, CBOR and other binary serialisation formats

For very large or frequently-updated fields, documenting a separate binary endpoint is essential so clients do not try to fetch kilobytes of unnecessary data. Debugging APIs with Base64 fields is straightforward with the right tool. Base64 encoder & decoder lets you decode any field locally in the browser, without storing it or sending it anywhere. Copy a Base64 field from a JSON response, paste it into the decoder and hit Decode. For text-like data (JSON inside the Base64, for example), the decoded output appears instantly.

For binary data like images, the hex view shows you the bytes. This helps confirm that the server sent what you expected and that your client decoder is working correctly. If a field decodes to unexpected data, the problem is in the server encoding or in how you are copying the field. If it decodes to a partial blob, the field may have been truncated or the Base64 length may be wrong. Local decoding speeds up debugging compared to writing the field to a file and opening external tools.

Takeaway: Base64 in JSON is a compromise, so document it — how the Base64 encoder & decoder helps you inspect and verify encoded fields locally

The practical approach to Base64 in APIs is awareness rather than avoidance. Base64 is the standard way to carry binary data in JSON and it works. Understand that each Base64 field costs a third more in size and a few milliseconds of CPU time per request-response cycle. For small critical metadata like authentication tokens (where the JWT itself is Base64-encoded), the cost is negligible. For large attachments, question whether the binary should travel in the same response or as a separate resource.

Document the encoding scheme and maximum sizes in your API specification. When inspecting responses, use Base64 encoder & decoder to verify the fields decode correctly and to understand what the server actually sent. That discipline keeps the trade-off visible and the decision deliberate rather than accidental.