Text to Binary

Convert text into the raw binary bytes behind each character.

Convert the letter A to binary and you get a tidy eight digits: 01000001. Convert an emoji the same way and the output is four times longer, because UTF-8 was never going to fit the whole world’s writing systems into the seven bits ASCII was built with.

One character, eight bits

The letter A has ASCII code 65. In binary, that is 01000001 — working from the right, the bits are worth 1, 2, 4, 8, 16, 32, 64 and 128, and 64 plus 1 equals 65. Text stored on a computer is simply a long run of these eight-bit groups, one per character.

Lowercase a, at 97, is 01100001 — the same pattern as A with one extra bit set, the bit worth 32, which is the entire mechanism behind the fixed 32-position gap between every uppercase and lowercase letter pair.

ASCII is a subset of UTF-8, until it is not

UTF-8 was deliberately designed so its first 128 characters are byte-for-byte identical to ASCII — a single byte, leading bit always 0. That is why plain English text looks exactly the same whether a tool treats it as ASCII or UTF-8.

Past code 127, UTF-8 switches to multi-byte sequences with specific bit patterns baked in: a leading byte starting 110 signals a two-byte character, 1110 signals three bytes, and 11110 signals four. Every continuation byte after the first starts with 10, which is how software can tell a continuation byte from the start of a new character.

A worked multi-byte example

Take the letter é (e with an acute accent), Unicode code point 233. It does not fit in one byte, so UTF-8 encodes it as two: 11000011 10101001, written in hex as C3 A9. The leading 110 marks a two-byte sequence, and the following 10 marks the continuation byte.

An emoji sits even further out, often needing all four bytes UTF-8 allows, with a leading byte starting 11110. That is the direct answer to why an emoji ‘weighs’ so much more than a letter when you are counting bytes.

Why an ASCII-only conversion mangles emoji and accents

  • ASCII has no codes above 127, so there is nothing for it to map an accented letter or an emoji onto.
  • A tool that reads text one byte at a time without knowing it is UTF-8 will split a multi-byte character apart and show each half as a separate, meaningless symbol.
  • This is the mechanical cause of the classic mojibake you sometimes see — readable text turning into strings of question marks or odd symbols after passing through the wrong encoding step.

Text to binary questions

How many bits does a single character use?

A plain ASCII character uses eight bits, one byte. A UTF-8 character outside the ASCII range can use two, three or four bytes, so 16, 24 or 32 bits, depending on which Unicode block it belongs to.

Why is the letter A specifically 01000001?

Because A was assigned decimal code 65 in the original 1963 ASCII standard, and 65 in binary is 01000001 — 64 plus 1, using the bit values 1, 2, 4, 8, 16, 32, 64 and 128.

Does converting text to binary lose any information?

No, as long as the encoding is known and consistent. Binary is simply a different way of writing the exact same byte values a computer already stores the text as.

Why do some binary conversions show 7 bits and others show 8?

Original ASCII only needed seven bits for its 128 codes, but computers store data in 8-bit bytes, so a leading zero is normally added to pad every ASCII character out to a full byte.

Can I convert binary back to the original text?

Yes, by splitting the binary into 8-bit groups, converting each to its decimal byte value, and looking up the corresponding character — provided the encoding used for multi-byte characters is known.

Cookie
We care about your data and would love to use cookies to improve your experience.