Binary translator
Text becomes binary by encoding each character as a number and writing that number in base two. This translator uses UTF-8, so a plain letter such as A is one byte; 01000001, while an accented character or an emoji takes two to four bytes. Decoding reverses the process, reading eight bits at a time.
How to translate binary
Most binary you find online is ASCII, seven bits per character padded to eight, which is why it decodes cleanly here. Anything beyond the first 128 characters is where encodings diverge: 01000011 01100011 is fine, but an é can be one byte in Latin-1 and two in UTF-8. This tool writes UTF-8 and reads it, falling back to raw bytes when a stream is not valid UTF-8 rather than refusing it.
The ASCII landmarks explain most of what you will see. Capital A is 65 and lowercase a is 97: exactly 32 apart, and 32 is a single bit, which is why 01000001 and 01100001 differ in one place and why toggling case used to be a bitwise operation rather than a lookup. Space is 32, digit zero is 48, and the printable range runs from 32 to 126. Everything below 32 is a control code, which is why a Windows line ending pasted in from a text file arrives as two bytes, 00001101 00001010, where a Unix one is a single 00001010.
Decoding is the fussier direction, because the bits carry no punctuation of their own and the split has to be assumed. This reads eight at a time, and falls back to seven when the total divides by seven and not by eight; a lot of pasted puzzle binary is seven-bit ASCII written without padding. Groups of mixed width cannot be recovered by any decoder, which is the whole argument for padding every character to eight bits when you write binary out in the first place.
What people use it for
- Decoding a binary string from a puzzle or a CTF
- Checking what bits a particular string produces
- Reading a bit-level capture back as text
- Checking a conversion you worked out by hand
- Demonstrating character encoding in a lesson
- Seeing how many bytes an accent or an emoji really costs
Questions
01000001, which is code 65 padded to eight bits. Lowercase a is 01100001.
Capital H, code 72. 01001000 01101001 is "Hi".
Yes. Anything that is not a 0 or a 1 is ignored, so line breaks, spaces and stray commas all work.
Eight, or seven throughout. The separators are ignored and the stream is split at a fixed width, so groups of mixed length will not decode.
It is handled. If the total is not a whole number of bytes but is a multiple of seven, it is read as seven-bit ASCII.
So that a decoder can split the stream at fixed intervals. Without padding, nothing in the bits says where one character ends and the next begins.
They are 32 apart, and 32 is a single bit: 01000001 against 01100001. Flipping it changes the case.
A space is 00100000, code 32. It encodes like any other character; the separators between groups are not part of the message.
Because UTF-8 is variable-length. Characters outside the basic Latin range take two, three or four bytes, and most emoji take four.
Both, and they agree for the first 128 characters. Above that this writes UTF-8, so é is 11000011 10101001 rather than a single byte.
It counts UTF-16 code units, and an emoji is two of them. Slicing at position 1 splits the pair and leaves a replacement diamond.
A MySQL utf8 column stores three bytes per character and emoji need four. The column type has to be utf8mb4.
They are read as single bytes instead of being refused, so a Latin-1 stream still comes back as something you can recognise rather than an error.
Not here; this writes base two. The ASCII table lists the hex code for every character below 128, and the number base converter moves a value between bases.
No. The conversion happens in this page, which is why it stays instant on a very large paste.