Computers use special codes for writing.
Computers use special codes for writing. 
Computers use codes to show text. In the past, these codes were hard to use together. One computer might see text as junk. 
Have you ever wondered how your phone knows exactly which letter to show?
Unicode works by giving every character a unique number. This number is called a code point. When you type a letter, the computer looks up its specific number. To store this data, the computer uses a method called encoding. Encoding is the way it turns those abstract numbers into binary data. There are different ways to do this, such as UTF-16 or UTF-32. However, UTF-8 is the most common encoding used by a large margin. It is very popular because it works well with older computer systems. This simple step allows text to flow smoothly across different devices.

Unicode is huge and keeps growing every year. Version 17.0 includes 159,801 different characters. It supports 172 different scripts, which are various ways of writing. The system is capable of holding more than 1.1 million characters in total. This massive space allows it to include historic writing like Egyptian hieroglyphs. It even includes 3,790 emoji, which are the small pictures we use in messages. These emoji became popular all over the world thanks to Unicode. The standard is also kept in sync with a system called ISO/IEC 10646.
Unicode is a universal character encoding standard used to digitize text from all the world's writing systems.
At its most abstract level, Unicode works by assigning a unique number to every character. This unique number is known as a code point. The standard focuses on encoding the underlying characters, which are called graphemes. It does not focus on the specific visual style of a character, such as its size or shape. Instead, the software used to view the text, like a web browser, decides how to render those styles. To store and process these code points, computers use a method called encoding. Encoding translates the abstract code points into sequences of binary data, which are bits and bytes.
Unicode is a massive and expanding system. Version 17.0 defines 159,801 characters and 172 different scripts. These scripts include various alphabets, abugidas, and syllabaries used in ordinary, literary, and technical contexts. The entire Unicode codespace is capable of encoding more than 1.1 million characters. This vast capacity allows the standard to include many historic scripts, such as Egyptian hieroglyphs. 
The history of Unicode began in the 1980s. A group of individuals with ties to Xerox's Character Code Standard began investigating a universal set. In 1987, Xerox employee Joe Becker started this work with Apple employees Lee Collins and Mark Davis. In August 1988, Becker published a draft proposal called "Unicode 88." He chose the name to suggest a unique, unified, and universal encoding. His original design used 16-bit characters to cover the world's living languages. He believed 16 bits would be enough to cover all modern-use characters. By 1991, the Unicode Consortium was incorporated in California. The first volume of the standard was published in October of that year.
As technology advanced, the system had to grow. In 1996, Unicode 2.0 implemented a surrogate character mechanism. This meant Unicode was no longer restricted to just 16 bits. This change increased the available space to over a million code points. This expansion allowed for the inclusion of thousands of rare or obsolete characters. It also allowed for the inclusion of complex characters used in names within CJK languages. To help with the transition from older systems, the first 256 code points mirror the ISO/IEC 8859-1 standard. This makes it easier to convert Western European text into Unicode without losing information.
The Unicode Consortium is a non-profit organization that coordinates all development. Its full members include major computer companies like Adobe, Apple, Google, IBM, Meta, Microsoft, Netflix, and SAP. The Consortium aims to replace all other encoding schemes with Unicode and its UTF schemes. They also manage the process for adding new scripts to the standard. Some scripts, like Jurchen, are currently on the Unicode Roadmap as candidates for encoding. Others, like the Rongorongo script, are waiting for user communities to provide proposals. Organizations like the Script Encoding Initiative also help fund the research needed for these additions.
Unicode provides much more than just a list of numbers. The standard includes detailed charts, reference data, and annexes for developers. These annexes explain complex concepts like character normalization and composition. They also cover character decomposition, collation, and directionality. This guidance is essential for programmers who need to implement text correctly. Even small details, like how to handle different accent marks, are documented.
🖼️ Images & Media (4)
More to explore
✨ What else?
Related topics you might enjoy
🔬 Go deeper
More advanced topics to explore
What is Nepedia?
A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.