One byte per letter, until the letter is not English
A computer stores text as bytes, and a byte is eight binary digits. Translating text to binary means writing out those bytes as 0s and 1s, so the only real question is which bytes a character becomes. For plain English the answer has been settled since ASCII: H is 01001000, a space is 00100000, and every translator agrees.
Beyond that, translators disagree. A byte only has 256 possible values, and there are far more characters than that. This page uses UTF-8, the encoding web pages, JSON and most modern software use, which spends one to four bytes on each character:
| Characters | Example | Bytes | Bit pattern |
|---|---|---|---|
| English letters, digits, basic punctuation (ASCII) | H | 1 | 0xxxxxxx |
| Accented Latin, Greek, Cyrillic, Hebrew, Arabic | é | 2 | 110xxxxx 10xxxxxx |
| Most Chinese, Japanese and Korean, the euro sign | € | 3 | 1110xxxx 10xxxxxx 10xxxxxx |
| Emoji and rare historic scripts | 👋 | 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
The pattern column is the clever part. The first bits of a byte say how many bytes the character has, and every byte after the first starts with 10, so a program can find where each character begins even in the middle of a stream. It is also why this page can tell you exactly which byte is wrong when a pasted message will not decode. Open Byte by byte under the tool to see the split for your own text.
Space, none or new line, and when you want numbers instead
- Space: one byte per group, the form puzzles, classroom exercises and most other translators use. Pick it unless you have a reason not to.
- None: a continuous stream of bits, for pasting into code or a field that counts every space against a limit. It is harder to check by eye.
- New line: one byte per line, which pastes into a spreadsheet as one byte per row and is easiest to annotate by hand.
The choice only affects what this page writes. Reading is the same for all three, plus two forms other tools produce: bytes split into four digit halves, and bytes with their leading zeros dropped, as long as a space separates them.
If what you have is a number rather than text, this is the wrong tool. Typed here, 42 is two characters and becomes 00110100 00110010. As a quantity it is 101010, which is what the number base converter gives you, along with hex, octal and two’s complement.
How to decode a binary message someone sent you
- Paste it into the Binary box as it is. Spaces, tabs and line breaks are ignored, so there is nothing to tidy first.
- Read the Text box. It updates on every keystroke, so you can also fix the binary in place and watch the text change.
- If a message appears instead, it names the problem. A character that is not a digit is given with its position, often a letter O typed for a zero or a smart quote picked up in a chat app.
- A count that does not divide by eight means a digit was lost or doubled. Look for a group that is shorter or longer than its neighbours.
- A byte that is not valid UTF-8 usually means the binary was made with another encoding, such as Latin-1, or that two digits were swapped inside a multi-byte character.
How to work out a character’s binary by hand
For an English letter, take its ASCII number and write it in eight binary places worth 128, 64, 32, 16, 8, 4, 2 and 1. H is 72, which is 64 plus 8, so the 64 and 8 places get a 1: 01001000.
Anything above 127 goes through the patterns in the table. Take é, code point U+00E9, which is 233 or 11101001 in binary:
- It needs more than seven bits, so it takes the two byte pattern, which has room for eleven. Pad it to eleven: 00011101001.
- Split it into five bits and six bits: 00011 and 101001.
- Put 110 in front of the first part and 10 in front of the second: 11000011 10101001.
That is exactly what the tool shows for é. Three and four byte characters work the same way with four and three bits in the first byte and six in each of the others.
Getting the binary where it needs to go
- Copy buttons: one under each box. Copy binary takes the separator you chose, so change it before copying, not after pasting.
- Length: every byte becomes eight characters, nine with a separator, so binary is roughly nine times longer than English and far more for emoji. Check the byte count under the boxes before pasting into anything with a limit.
- Spreadsheets: New line puts one byte in each row. Format the column as text first, or the software may read 00100000 as the number 100000 and drop the zeros.
- If copying fails: in-app browsers inside social apps sometimes refuse clipboard access, and a warning says so. Select the box and copy by hand, or open the page in your normal browser.
- Nothing is in the address bar, so sharing the link shares the tool, not your message.
The same conversion in Python, JavaScript, Excel and a shell
- Python:
' '.join(f'{b:08b}' for b in text.encode())encodes, andbytes(int(b, 2) for b in s.split()).decode()decodes. Both default to UTF-8. - JavaScript:
new TextEncoder()gives the UTF-8 bytes, which is what this page uses. AvoidcharCodeAt(0).toString(2), which returns a UTF-16 code unit, not bytes: é comes out as 11101001, and an emoji as half of a surrogate pair. - Excel and Google Sheets:
=DEC2BIN(CODE(A1), 8)works for plain ASCII letters only. CODE does not return UTF-8 bytes, and DEC2BIN stops at 511. - Linux and macOS terminals:
printf 'Hi' | xxd -bprints each byte in binary, next to an offset and a text column. Use printf rather than echo, which adds a line break byte at the end.
Why UTF-8 and not ASCII or UTF-16
ASCII covers 128 characters and stops. UTF-16, used inside JavaScript, Java and Windows, stores most characters in two bytes, so its binary for English has an extra byte of zeros per letter. UTF-8 matches ASCII byte for byte for English and still handles every other script, which is why the web settled on it, and why the WHATWG Encoding Standard tells new formats to use it.
When two strings that look the same have different binary
- Accents can be stored two ways. é may be one code point, 11000011 10101001, or a plain e followed by a combining accent, 01100101 11001100 10000001. They look the same, compare as different and give different binary. Byte by byte shows which one you have.
- Line endings differ. Text from Windows often ends lines with 00001101 00001010, carriage return plus line feed, where macOS and Linux use only 00001010.
- Invisible characters count. A byte order mark (11101111 10111011 10111111) at the start of a file, or a zero width space picked up from a web page, adds bytes you cannot see. They appear by name in Byte by byte.
- Binary is not a secret. Anyone can decode it. Never send a password or a key in binary on the assumption nobody will bother; use the hash generator if you need a one way fingerprint instead.
- Other encodings will not match. Binary produced from Latin-1, Windows-1252 or UTF-16 is correct for that encoding but reads as an error here, or as different characters, for anything outside plain English.
Frequently asked questions
How do I translate binary to text?
Paste the binary into the Binary box and the Text box fills in as you type. Spaces, tabs and line breaks are ignored, so bytes separated by spaces, one per line or run together all work. If something cannot be read, the message under the separator buttons says which digit or byte is wrong, and the Text box keeps the last result that did decode instead of going blank.
What does 01001000 01101001 mean?
It spells Hi. The first byte, 01001000, is 72 in decimal, which is the capital H in ASCII and in UTF-8. The second, 01101001, is 105, a lowercase i. Plain English letters, digits and punctuation take one byte each, so a message in English is eight binary digits per character, spaces included.
How do you write I love you in binary?
In UTF-8, and in plain ASCII, which gives the same bytes for these letters, it is 01001001 00100000 01101100 01101111 01110110 01100101 00100000 01111001 01101111 01110101. That is ten bytes, eight letters and two spaces. The space is 00100000, which is why the same byte shows up twice.
Why is an emoji 32 bits and not 8?
One byte holds 256 values, which is not enough for every character, so UTF-8 uses between one and four bytes per code point. English letters take one, accented letters such as the e in cafe with an acute take two, most Chinese, Japanese and Korean characters take three, and emoji take four. Many emoji are several code points glued together: a thumbs up with a skin tone is two code points and eight bytes, and a family of four joined with invisible joiners is 25 bytes.
Why does another binary translator give a different answer?
For plain English letters every translator agrees, because ASCII and UTF-8 give the same bytes. The answers split on anything else. Some tools write out the character number instead of its UTF-8 bytes, so an accented e comes out as 11101001, its Latin-1 value, rather than the two UTF-8 bytes 11000011 10101001. Others drop the leading zeros and write H as 1001000. Neither is wrong for its own purpose, but only UTF-8 bytes decode correctly in a modern program.
Why does it say the binary is not a whole number of bytes?
Each byte is exactly eight digits, so the total number of 0s and 1s has to divide by eight. When it does not, a digit was lost or doubled while copying, and the message tells you how many bits are left over. If your binary came from a tool that drops leading zeros, keep a space between the bytes: groups of six or seven digits are then read as one byte each.
Can I decode binary that has no spaces?
Yes, as long as it has not also lost its leading zeros. A continuous run of digits is cut into bytes every eight digits from the left, which is only right when every byte kept all eight. Choose None under Separate bytes with if you want to produce that form yourself; the other two options put a space or a line break between bytes.
Is writing a message in binary a way to hide it?
No. Binary is a spelling of the bytes, not encryption, and anyone can turn it back into text in seconds with a page like this one. It is fine for puzzles, classroom exercises and tattoos, but anything that has to stay private needs real encryption, and a password should never be sent in binary or in Base64 on the assumption that nobody will bother to decode it.
What is the difference between this and a binary number converter?
This page converts characters. The text 42 is two characters, so it becomes two bytes, 00110100 00110010. A number converter treats 42 as a quantity and gives 101010. If you want the binary value of a number, or to convert between binary, decimal and hex, use the number base converter linked below the article instead.
Is my text sent anywhere?
No. Encoding and decoding run in your browser with the TextEncoder and TextDecoder functions every modern browser has, and nothing you type is uploaded, put in the address bar or saved between visits. The only thing recorded is an anonymous event saying that a copy button was pressed and whether it copied the text or the binary.
Sources
- IETF, RFC 3629: UTF-8, a transformation format of ISO 10646, for the byte patterns and the byte values UTF-8 forbids.
- WHATWG, Encoding Standard, which defines the UTF-8 decoder browsers implement, including how it rejects overlong forms and surrogates.
- MDN Web Docs, TextEncoder and TextDecoder, the two browser functions this page uses, and the fatal and ignoreBOM options that keep a round trip exact.
- The code point to byte table is the well-formed byte sequence table in chapter 3 of The Unicode Standard, which RFC 3629 and the WHATWG decoder both follow.
More encoding tools
For numbers rather than text, the number base converter converts between binary, decimal, octal and hex. To carry bytes through email or JSON, the Base64 encoder is shorter than binary, and the URL encoder shows the same UTF-8 bytes as percent signs and hex. The word counter also reports UTF-8 bytes for a whole document.
Related tools
- Number Base Converter
Convert numbers between binary, octal, decimal and hexadecimal, including large integers a…
- Base64 Encode and Decode
Convert text to Base64 and back with full UTF-8 support, so accented letters, CJK and emoj…
- URL Encoder and Decoder
Percent-encode or decode URLs and query strings with encodeURI or encodeURIComponent semantics.
- SHA-256 Hash Generator
Compute SHA-1, SHA-256, SHA-384 and SHA-512 hashes of text or a file with the browser's Web Crypto API.
Privacy: encoding and decoding run entirely in your browser. What you type in either box is never uploaded, never written into the address bar and never stored between visits. Pressing a copy button records an anonymous analytics event naming only whether the text or the binary was copied, never the content.
