# Why just UTF-16? Add UTF-8 support everywhere!

**URL:** <https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177>\
**Category:** 💡 Ideas\
**Created:** [January 26, 2022, 5:12am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177 "2022-01-26T05:12:18Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![d1gital\_love](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@d1gital\_love](https://es.discourse.group/u/d1gital_love)\
**Post date:** [January 26, 2022, 5:12am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/1 "2022-01-26T05:12:18Z")

</div>

UTF-8 is the most common encoding on the web.

How ECMAScript® 2022 Language Specification (January 20, 2022) just ignores UTF-8?  
[https://tc39.es/ecma262/multipage/ecmascript-data-types-and-values.html](https://tc39.es/ecma262/multipage/ecmascript-data-types-and-values.html)  
[https://tc39.es/ecma262/multipage/text-processing.html](https://tc39.es/ecma262/multipage/text-processing.html)

I don't know how old text with mandatory UTF-16 is, but it's just awful and not acceptable.

This may lead to many minor changes almost everywhere where sequences or characters can be.

---

<div class="post-metadata">

**Author:** ![swhiteman](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/swhiteman/32/716_2.png) [@swhiteman](https://es.discourse.group/u/swhiteman)\
**Post date:** [January 26, 2022, 7:55am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/2 "2022-01-26T07:55:13Z")

</div>

Hi there,

Just like Java and the .NET CLR (and countless other still-modern environments), ES _internally_ represents a String as a sequence of UTF-16 code units. You may be confusing that _internal_ representation with UTF-8 encoding of _text output_ or decoding of _text input_. But those are different concerns: a String is not the same as "text," despite the terms being used interchangeably at times.

The ES spec itself is agnostic about text encodings used at runtime; those are governed by other standards.

For example, in the browser, you would use the standard [TextEncoder/TextDecoder](https://encoding.spec.whatwg.org/) which strongly emphasizes UTF-8 and is also implemented in NodeJS. When you use the [Fetch API](https://fetch.spec.whatwg.org/) the encoding is contained in the `charset` header (which defaults to UTF-8). Ditto the JSON standard. Note Encoding and Fetch are standardized _JavaScript APIs_, but they are separate from the language itself.

Perhaps you could explain more your proposal to have UTF-8 be used as the internal representation of Strings? It's true that very modern languages (Go, Rust) use UTF-8 internally and there's a slow creep  
— enabled of course by faster processors and cheaper storage — toward UTF-8 in brand new technologies. But establishing UTF-16 to be "awful and not acceptable" is a tall order, given its continued dominance.

---

<div class="post-metadata">

**Author:** ![tabatkins](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/tabatkins/32/170_2.png) [@tabatkins](https://es.discourse.group/u/tabatkins)\
**Post date:** [January 26, 2022, 11:16pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/3 "2022-01-26T23:16:29Z")

</div>

There is, unfortunately, absolutely zero chance of changing the way that JS strings encode their characters. That would be a massive breaking change across the entire ecosystem.

Modern string APIs often address text as codepoints, like `String.prototype.codePointAt()`, which is generally what you want.

(You rarely, if ever, actually want UTF-8; encoding details are a detail you rarely want to be aware of, instead of just getting codepoints and/or grapheme clusters. The problem with JS is that it exposes encoding details, and in particular it exposes details of a _really bad_ encoding (UCS-2-ish, not even UTF-16).)

---

<div class="post-metadata">

**Author:** ![d1gital\_love](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@d1gital\_love](https://es.discourse.group/u/d1gital_love)\
**Post date:** [January 27, 2022, 6:13am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/4 "2022-01-27T06:13:42Z")

</div>

> [@tabatkins](#):
>
> You rarely, if ever, actually want UTF-8;

I think we need UTF-8 support badly. UTF-8 can support all of Unicode. UTF-8 is just old ASCII in some sense. Many good programs use UTF-8 by default. ECMAScript and JavaScript are almost completely going against the flow.

I don't use Windows much ATM but [some say they still had ANSI in 2010](https://answers.microsoft.com/en-us/windows/forum/windows_7-windows_programs/default-utf-8-encoding-for-new-notepad-documents/525f0ae7-121e-4eac-a6c2-cfe6b498712c).

---

<div class="post-metadata">

**Author:** ![aclaymore](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/aclaymore/32/501_2.png) [@aclaymore](https://es.discourse.group/u/aclaymore)\
**Post date:** [January 27, 2022, 10:27am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/5 "2022-01-27T10:27:56Z")

</div>

@d1gital_love did you have a particular use case.

While EcmaScript itself may not reference UTF-8. Many JS platforms do have APIs that support other encodings.

> **[TextEncoder - Web APIs | MDN](https://developer.mozilla.org/en-US/docs/Web/API/TextEncoder)**
>
> The TextEncoder interface takes a stream of code points as input and emits a stream of UTF-8 bytes.

[https://nodejs.org/api/string\_decoder.html](https://nodejs.org/api/string_decoder.html)

> **[TextEncoder and TextDecoder in Deno](https://medium.com/deno-the-complete-reference/textencoder-and-textdecoder-in-deno-cfca83be1792)**
>
> Convert raw data to string and vice versa using TextDecoder and TextEncoder

---

<div class="post-metadata">

**Author:** ![d1gital\_love](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@d1gital\_love](https://es.discourse.group/u/d1gital_love)\
**Post date:** [January 27, 2022, 10:48am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/6 "2022-01-27T10:48:49Z")

</div>

I don't think that many know about TextEncoder and TextDecoder but many know about String type.

Again:

> String.prototype.codePointAt (pos)  
> Returns a non-negative [integral Number](https://tc39.es/ecma262/multipage/notational-conventions.html#integral-number) less than or equal to 0x10FFFF𝔽 that is the numeric value of the UTF-16 encoded code point

[https://tc39.es/ecma262/multipage/text-processing.html#sec-string.prototype.codepointat](https://tc39.es/ecma262/multipage/text-processing.html#sec-string.prototype.codepointat)

> **[String.prototype.codePointAt() - JavaScript | MDN](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/String/codePointAt)**
>
> The codePointAt() method of String values returns a non-negative integer that is the Unicode code point value of the character starting at the given index. Note that the index is still based on UTF-16 code units, not Unicode code points.

---

<div class="post-metadata">

**Author:** ![theScottyJam](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/thescottyjam/32/616_2.png) [@theScottyJam](https://es.discourse.group/u/theScottyJam)\
**Post date:** [January 27, 2022, 2:47pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/7 "2022-01-27T14:47:14Z")

</div>

Alright, I think @tabatkins is agreeing with you that this is a real issue, that JavaScript's default encoding behavior isn't the greatest, but he also shared an important point - it's not like we can just _change_ JavaScript strings from UTF-16 to UTF-8 without breaking old websites.

So, do you have a solution to propose of how we should go about adding support for UTF-8 without breaking older JavaScript? If so, please share, and we can have a discussion around it.

---

<div class="post-metadata">

**Author:** ![tabatkins](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/tabatkins/32/170_2.png) [@tabatkins](https://es.discourse.group/u/tabatkins)\
**Post date:** [January 27, 2022, 5:54pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/8 "2022-01-27T17:54:14Z")

</div>

Correct.

Again, UTF-8 is not "Unicode", it's an _encoding_ of Unicode; a way of turning unicode characters into bits (and back). JS already supports all of Unicode. The default string indexing (`"foo"[0]`) is busted, because it indexes the string according to UCS-2 code units, rather than characters. That's unfortunate and bad, but it's impossible to change. JS has many new ways of interacting with strings that _do_ work on characters - `[..."foo"]` is character-based, `"foo".codePointAt(0)` is character-based, `String.fromCodePoint(0xfffd)` is character-based. Regexes also have recently gained ways of interacting with strings properly as Unicode characters (and are continuing to evolve in that direction).

So everything you need is already present or upcoming. We're stuck with the bad parts forever.

---

<div class="post-metadata">

**Author:** ![d1gital\_love](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@d1gital\_love](https://es.discourse.group/u/d1gital_love)\
**Post date:** [January 27, 2022, 7:23pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/9 "2022-01-27T19:23:39Z")

</div>

If we take `String.prototype.codePointAt(offset)` as example then we can add `outputEncoding` argument with default parameters with old (`UTF-16`) encoding to make it backward-compatible.

[https://tc39.es/ecma262/multipage/ecmascript-language-functions-and-classes.html#sec-function-definitions](https://tc39.es/ecma262/multipage/ecmascript-language-functions-and-classes.html#sec-function-definitions)

---

<div class="post-metadata">

**Author:** ![tabatkins](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/tabatkins/32/170_2.png) [@tabatkins](https://es.discourse.group/u/tabatkins)\
**Post date:** [January 28, 2022, 11:51pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/10 "2022-01-28T23:51:56Z")

</div>

Again, that would not do anything like what you want. A string is not an encoding. The TextEncoder API, which outputs a TypedArray of encoded binary data, can output a string encoded as UTF-8.

---

<div class="post-metadata">

**Author:** ![theScottyJam](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/thescottyjam/32/616_2.png) [@theScottyJam](https://es.discourse.group/u/theScottyJam)\
**Post date:** [January 29, 2022, 12:13am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/11 "2022-01-29T00:13:49Z")

</div>

It looks like that MDN snippet you linked to was recently updated in the last few days to say "unicode" instead of "UTF-16". From what I understand, (git history says it happened 18 days ago, but I'm not sure how often MDN updates its content from the git repo, so perhaps you were viewing the older content?). From what I understand, `codePointAt()` returns a number that's assigned to a specific unicode character (like a unique id for that character), which is irrelevant to however the engine may be encoding that specific character.

---

<div class="post-metadata">

**Author:** ![mhofman](https://avatars.discourse-cdn.com/v4/letter/m/f14d63/32.png) [@mhofman](https://es.discourse.group/u/mhofman)\
**Post date:** [January 30, 2022, 4:16am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/12 "2022-01-30T04:16:47Z")

</div>

> [@tabatkins](#):
>
> `"foo".codePointAt(0)` is character-based

To be fair, `codePointAt` is not entirely character based since the position argument is still based on the UCS-2 representation.

---

<div class="post-metadata">

**Author:** ![d1gital\_love](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@d1gital\_love](https://es.discourse.group/u/d1gital_love)\
**Post date:** [January 30, 2022, 5:49am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/13 "2022-01-30T05:49:00Z")

</div>

> [@theScottyJam](#):
>
> From what I understand, `codePointAt()` returns a number that's assigned to a specific unicode character (like a unique id for that character), which is irrelevant to however the engine may be encoding that specific character.

Yes. `codePointAt` was pointless as example but not something like `charAt()`...

> The [`String`](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/String) object's **`charAt()`** method returns a new string consisting of the single UTF-16 code unit located at the specified offset into the string.

> **[String.prototype.charAt() - JavaScript | MDN](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/String/charAt)**
>
> The charAt() method of String values returns a new string consisting of the single UTF-16 code unit at the given index.

---

<div class="post-metadata">

**Author:** ![claudiameadows](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/claudiameadows/32/126_2.png) [@claudiameadows](https://es.discourse.group/u/claudiameadows)\
**Post date:** [February 2, 2022, 3:00am UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/14 "2022-02-02T03:00:55Z")

</div>

If you want models of working with characters:

- UTF-8 is the typical interchange format, but is as bad as UTF-16 when it comes to string _manipulation and processing_.
- UTF-32 solves the code point problem and [is what Python (as of 3.3) uses internally for strings that contain emojis and the like](https://legacy.python.org/dev/peps/pep-0393/), but requires a lot of memory to sustain (hence why Python tries to avoid it where it can) and doesn't account for multi-code point graphemes.
- [Swift addresses that very well by having strings be sequences of extended grapheme clusters with views for both UTF-8, UTF-16, and UTF-32](https://docs.swift.org/swift-book/LanguageGuide/StringsAndCharacters.html), but it comes at the cost of a mildly bloated API and a number of performance caveats (like most indexed access operations being `O(n)`, including string slicing) that make it very inefficient and suboptimal for structured data parsing.

---

<div class="post-metadata">

**Author:** ![Rudxain](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/rudxain/32/2272_2.png) [@Rudxain](https://es.discourse.group/u/Rudxain)\
**Post date:** [September 14, 2022, 10:30pm UTC](https://es.discourse.group/t/why-just-utf-16-add-utf-8-support-everywhere/1177/15 "2022-09-14T22:30:55Z")

</div>

[Related](https://hsivonen.fi/string-length)
