Text Encoding

Binary translator

Text 5 chars
Binary 44 chars
01001000 01100101 01101100 01101100 01101111
Live · UTF-8, eight bits per byte

Text becomes binary by encoding each character as a number and writing that number in base two. This translator uses UTF-8, so a plain letter such as A is one byte; 01000001, while an accented character or an emoji takes two to four bytes. Decoding reverses the process, reading eight bits at a time.

How to translate binary

1 Type or paste your text. The binary appears beside it as you type.
2 Press Swap to go the other way and paste binary in instead.
3 Drop a .txt file onto the box to load a longer piece of text.

Most binary you find online is ASCII, seven bits per character padded to eight, which is why it decodes cleanly here. Anything beyond the first 128 characters is where encodings diverge: 01000011 01100011 is fine, but an é can be one byte in Latin-1 and two in UTF-8. This tool writes UTF-8 and reads it, falling back to raw bytes when a stream is not valid UTF-8 rather than refusing it.

The ASCII landmarks explain most of what you will see. Capital A is 65 and lowercase a is 97: exactly 32 apart, and 32 is a single bit, which is why 01000001 and 01100001 differ in one place and why toggling case used to be a bitwise operation rather than a lookup. Space is 32, digit zero is 48, and the printable range runs from 32 to 126. Everything below 32 is a control code, which is why a Windows line ending pasted in from a text file arrives as two bytes, 00001101 00001010, where a Unix one is a single 00001010.

Decoding is the fussier direction, because the bits carry no punctuation of their own and the split has to be assumed. This reads eight at a time, and falls back to seven when the total divides by seven and not by eight; a lot of pasted puzzle binary is seven-bit ASCII written without padding. Groups of mixed width cannot be recovered by any decoder, which is the whole argument for padding every character to eight bits when you write binary out in the first place.

What people use it for

  • Decoding a binary string from a puzzle or a CTF
  • Checking what bits a particular string produces
  • Reading a bit-level capture back as text
  • Checking a conversion you worked out by hand
  • Demonstrating character encoding in a lesson
  • Seeing how many bytes an accent or an emoji really costs

Questions

01000001, which is code 65 padded to eight bits. Lowercase a is 01100001.

Unicode Consortium, UTF-8 and the encoding forms
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding