Base64 in JavaScript: The Parts That Bite
Base64 encoding is simple enough that most developers never read about it, which is why the same four bugs keep appearing. All of them come from the same root cause: Base64 encodes bytes, and JavaScript strings are not bytes.
1. btoa throws on anything outside Latin-1
btoa is a fossil from 1995. It expects a "binary string" โ one where every character code is 0โ255 โ and throws on anything else. JavaScript strings are UTF-16, so รฉ (U+00E9) happens to fit, but ๐ (U+1F389) does not, and neither does most of the world's writing.
The correct approach encodes to UTF-8 bytes first:
You will find unescape(encodeURIComponent(str)) recommended widely. It works, but unescape has been deprecated since ECMAScript 3 and is in the specification's Annex B. Use TextEncoder.
The Base64 encoder on this site does the TextEncoder conversion for you, so emoji, CJK text and accented characters round-trip correctly.
2. Base64 and base64url are different alphabets
Standard Base64 (RFC 4648 ยง4) uses + and / as its last two characters and pads with =. All three are hostile in a URL: + means a space in query strings, / is a path separator, and = separates parameters.
base64url (RFC 4648 ยง5) swaps them and usually drops padding:
| Standard | base64url | |
|---|---|---|
| Index 62 | + | - |
| Index 63 | / | _ |
| Padding | = required | usually omitted |
JWTs use base64url for all three segments. If you are decoding a token by hand and atob throws, this is why โ and it is also why a JWT that works in one library fails in another that assumes standard Base64.
3. Node's decoder accepts garbage silently
This is the one that reaches production. Buffer.from(str, "base64") does not validate:
Node skips characters outside the alphabet and decodes what remains. A corrupted payload therefore decodes to something plausible instead of failing, and the bug surfaces much later as mysterious truncation. If the input comes from outside your system, validate before decoding:
Node 16+ also offers Buffer.from(str, "base64url"), which handles the URL-safe alphabet and missing padding natively โ prefer it over hand-rolled replacement when you are on a recent runtime.
4. Padding rules, and when you can drop it
Base64 processes three bytes at a time into four characters. When the input length is not a multiple of three, the output is padded with = to a multiple of four:
| Input bytes | Output | Padding |
|---|---|---|
| 3, 6, 9 โฆ | 4, 8, 12 โฆ | none |
| 1 remaining | 2 characters | == |
| 2 remaining | 3 characters | = |
Padding carries no information โ the decoder can infer length from the character count. It exists for concatenated streams. You can drop it when both ends agree, which is exactly what base64url does, but a strict decoder will reject an unpadded string, so never drop it unilaterally.
A useful consequence: = can only ever appear at the end. A = in the middle of what should be Base64 means two payloads were concatenated, or the string was truncated and re-padded.
5. Size, and the thing Base64 is not
Base64 costs exactly 33% โ four characters per three bytes โ plus padding. A 3 MB image becomes a 4 MB data URL. That matters for inlining assets: below roughly 1โ2 KB the saved HTTP request usually wins, above that it usually does not, and a Base64 data URL cannot be cached separately from the document that contains it.
And the thing that has to be said: Base64 is not encryption, and it is not obfuscation. It is a reversible transport encoding with no key. A Base64-encoded password is a plaintext password with extra steps. HTTP Basic auth works exactly this way, which is precisely why it must never be used without TLS.
6. Whitespace and line breaks
MIME (RFC 2045) requires Base64 to be wrapped at 76 characters. PEM certificates wrap at 64. So real-world Base64 frequently arrives with embedded newlines, and decoders disagree about them:
| Decoder | Newlines in input |
|---|---|
Browser atob | Tolerates ASCII whitespace |
Node Buffer.from | Ignored, along with everything else invalid |
Python base64.b64decode | Throws unless validate=False |
Java Base64.getDecoder | Throws โ use getMimeDecoder instead |
If you are moving Base64 between languages, strip whitespace on the way in: input.replace(/\s/g, ""). It is never wrong and it removes an entire category of cross-platform bug.
Related
- Base64 encoder and decoder โ UTF-8 safe, runs in your browser
- Every JSON.parse error explained
- JSONPath tester