Log in Sign up
Back to Discover
💻

Unicode

technology Maturity 11-13 war conflict
This article covers sensitive topics: war_conflict. Parents can manage visibility in Parental Controls.

Computers use special codes for writing.

Unicode sample.svg
Unicode sample.svg
These codes help us use all kinds of letters. They even help us use small smiley faces. This helps people all over the world talk. It is very helpful. Can you use a smiley face?

43 words

Computers use special codes for writing.

Unicode sample.svg
Unicode sample.svg
These codes help us use all kinds of letters. They work for many ways of writing. This helps people all over the world talk.
Hiero O4.png
Hiero O4.png
The codes can even show old symbols. They can show small smiley faces too. These faces are called emoji.
Cyrillic cursive.svg
Cyrillic cursive.svg
Most of the text on the web uses these codes. It is a very big set of symbols. It helps computers understand everyone.

78 words

Computers use codes to show text. In the past, these codes were hard to use together. One computer might see text as junk.

Unicode sample.svg
Unicode sample.svg
Unicode is a big set of codes. It helps computers use all writing systems. This includes many ways of writing.
Cyrillic cursive.svg
Cyrillic cursive.svg
A group called the Unicode Consortium manages it. They work to keep the codes updated. Version 17.0 has 159,801 characters. It also supports 172 different scripts. These scripts are ways of writing.
Hiero O4.png
Hiero O4.png
Unicode can even show old symbols. For example, it can show Egyptian hieroglyphs. It also includes 3,790 emoji. These are the small pictures used in messages. Unicode can hold over 1.1 million characters. To use these codes, computers use encodings. An encoding is a way to turn codes into data. UTF-8 is the most common encoding used today. It helps many different systems work well together. This makes it easy to read text on the web.

157 words

Have you ever wondered how your phone knows exactly which letter to show?

Unicode sample.svg
Unicode sample.svg
Computers do not actually see letters or numbers like we do. Instead, they use a system of codes to represent every single character. This system is called Unicode. It is a global standard that helps computers use all the world's writing systems. Without Unicode, a message sent from one computer might look like total junk on another. It makes sure that text stays clear and correct everywhere on the internet.
Cyrillic cursive.svg
Cyrillic cursive.svg

Unicode works by giving every character a unique number. This number is called a code point. When you type a letter, the computer looks up its specific number. To store this data, the computer uses a method called encoding. Encoding is the way it turns those abstract numbers into binary data. There are different ways to do this, such as UTF-16 or UTF-32. However, UTF-8 is the most common encoding used by a large margin. It is very popular because it works well with older computer systems. This simple step allows text to flow smoothly across different devices.

Hiero O4.png
Hiero O4.png
The idea for Unicode started in the 1980s. A group of people wanted to fix the problem of incompatible text sets. In 1987, Joe Becker from Xerox began working on this idea. He worked with Lee Collins and Mark Davis from Apple. In August 1988, Becker published a draft proposal for the system. He chose the name Unicode to suggest a unique and universal way to encode text. By 1991, the Unicode Consortium was officially formed in California. They released the first volume of the standard in October of that year.

Unicode is huge and keeps growing every year. Version 17.0 includes 159,801 different characters. It supports 172 different scripts, which are various ways of writing. The system is capable of holding more than 1.1 million characters in total. This massive space allows it to include historic writing like Egyptian hieroglyphs. It even includes 3,790 emoji, which are the small pictures we use in messages. These emoji became popular all over the world thanks to Unicode. The standard is also kept in sync with a system called ISO/IEC 10646.

I acute - soft dotted and Lithuanian dot.svg
I acute - soft dotted and Lithuanian dot.svg
You can see Unicode in action every time you use a web browser. It is the reason you can read news from different countries easily. It links all the different ways people write into one single digital language. Whether you are reading a math equation or a simple text message, Unicode is working. It acts like a giant, invisible library for every symbol ever used. This helps developers build software that people everywhere can use. It truly connects the whole digital world through text.

458 words

Unicode is a universal character encoding standard used to digitize text from all the world's writing systems.

Unicode sample.svg
Unicode sample.svg
It acts as a common language for computers to represent letters, numbers, and symbols. Before Unicode, different computers used many incompatible character sets. These sets often only worked for specific regions or specific computer architectures. If you sent text from one system to another, it often appeared as unreadable garbage characters. Unicode solved this by merging these various sets into one single, massive repertoire. Today, it is used for the vast majority of text on the Internet. It is a critical consideration for any modern software development.

At its most abstract level, Unicode works by assigning a unique number to every character. This unique number is known as a code point. The standard focuses on encoding the underlying characters, which are called graphemes. It does not focus on the specific visual style of a character, such as its size or shape. Instead, the software used to view the text, like a web browser, decides how to render those styles. To store and process these code points, computers use a method called encoding. Encoding translates the abstract code points into sequences of binary data, which are bits and bytes.

Cyrillic cursive.svg
Cyrillic cursive.svg
The standard defines three main encodings: UTF-8, UTF-16, and UTF-32. UTF-8 is the most widely used encoding by a very large margin. This is partly because it is backwards-compatible with the older ASCII standard.

Unicode is a massive and expanding system. Version 17.0 defines 159,801 characters and 172 different scripts. These scripts include various alphabets, abugidas, and syllabaries used in ordinary, literary, and technical contexts. The entire Unicode codespace is capable of encoding more than 1.1 million characters. This vast capacity allows the standard to include many historic scripts, such as Egyptian hieroglyphs.

Hiero O4.png
Hiero O4.png
Beyond letters, Unicode also encodes 3,790 emoji. The widespread adoption of Unicode was a major reason emoji became popular outside of Japan. The Unicode repertoire is also kept synchronized with ISO/IEC 10646. These two standards are code-for-code identical with one another.

The history of Unicode began in the 1980s. A group of individuals with ties to Xerox's Character Code Standard began investigating a universal set. In 1987, Xerox employee Joe Becker started this work with Apple employees Lee Collins and Mark Davis. In August 1988, Becker published a draft proposal called "Unicode 88." He chose the name to suggest a unique, unified, and universal encoding. His original design used 16-bit characters to cover the world's living languages. He believed 16 bits would be enough to cover all modern-use characters. By 1991, the Unicode Consortium was incorporated in California. The first volume of the standard was published in October of that year.

As technology advanced, the system had to grow. In 1996, Unicode 2.0 implemented a surrogate character mechanism. This meant Unicode was no longer restricted to just 16 bits. This change increased the available space to over a million code points. This expansion allowed for the inclusion of thousands of rare or obsolete characters. It also allowed for the inclusion of complex characters used in names within CJK languages. To help with the transition from older systems, the first 256 code points mirror the ISO/IEC 8859-1 standard. This makes it easier to convert Western European text into Unicode without losing information.

The Unicode Consortium is a non-profit organization that coordinates all development. Its full members include major computer companies like Adobe, Apple, Google, IBM, Meta, Microsoft, Netflix, and SAP. The Consortium aims to replace all other encoding schemes with Unicode and its UTF schemes. They also manage the process for adding new scripts to the standard. Some scripts, like Jurchen, are currently on the Unicode Roadmap as candidates for encoding. Others, like the Rongorongo script, are waiting for user communities to provide proposals. Organizations like the Script Encoding Initiative also help fund the research needed for these additions.

Unicode provides much more than just a list of numbers. The standard includes detailed charts, reference data, and annexes for developers. These annexes explain complex concepts like character normalization and composition. They also cover character decomposition, collation, and directionality. This guidance is essential for programmers who need to implement text correctly. Even small details, like how to handle different accent marks, are documented.

I acute - soft dotted and Lithuanian dot.svg
I acute - soft dotted and Lithuanian dot.svg
By providing these rules, Unicode ensures that text remains consistent across every platform in the world.

739 words
🖼️ Images & Media (4)
File:Unicode sample.svg
Unicode sample.svg
File:Hiero O4.png
Hiero O4.png
File:Cyrillic cursive.svg
Cyrillic cursive.svg
File:I acute - soft dotted and Lithuanian dot.svg
I acute - soft dotted and Lithuanian dot.svg
Up Next
💻
UTF-8
Technology
More to explore

🪜 Step back

Simpler topics to build understanding

What is Nepedia?

A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.