# Code point string

**URL:** <https://es.discourse.group/t/code-point-string/1351>\
**Category:** 💡 Ideas\
**Created:** [May 25, 2022, 3:05pm UTC](https://es.discourse.group/t/code-point-string/1351 "2022-05-25T15:05:28Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Klaider](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/klaider/32/1447_2.png) [@Klaider](https://es.discourse.group/u/Klaider)\
**Post date:** [May 25, 2022, 3:05pm UTC](https://es.discourse.group/t/code-point-string/1351/1 "2022-05-25T15:05:28Z")

</div>

In Python, the string data type consists of Unicode Scalar Values, or code points in short, while in ECMAScript the string data type consists of Unicode Code Units. In ECMAScript 4 reference interpreter, I saw the string data type consisting of code points, so I got this idea from there, to start with.

Python runtime uses a technique like this to encode strings: if there's no code point that is ≥ U+100, the string data type is encoded using 1-byte per character. If there's any code point that is ≥ U+100, but no code point ≥ U+10000, it's encoded using 2-byte per character; otherwise it uses 4-byte per character.

Since ECMAScript uses the `number` data type, whichs supports `0x10FFFF` value, it's possible to have Python-based string data type by having alternate versions of the string manipulation methods. By alternate versions I mean, there can be an option, like `'use code point'`, which will cause manipulations to be code-point-based, and not code-unit-based.

So the idea is not to add code-point-specific methods, but to allow existing methods to work either for code units or for code points.

So, for example, it'd work like this:

```javascript
// actual behavior
'\u{10ffff}'.charCodeAt(0); // U+DBFF
'\u{10ffff}'.charCodeAt(1); // U+DFFF

// desired behavior
'use code point';
'\u{10ffff}'.charCodeAt(0); // U+10FFFF
'\u{10ffff}'.charCodeAt(1); // NaN

```

Some compilers could automatically add this `'use code point'` directive. This could even be set at HTML-`<script>`-level.

#### Implementation

To implement this feature in V8 is required a four-byte representation of the `string` type and support 0-`0x10ffff` range for character codes.

---

<div class="post-metadata">

**Author:** ![DiriectorDoc](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/diriectordoc/32/987_2.png) [@DiriectorDoc](https://es.discourse.group/u/DiriectorDoc)\
**Post date:** [May 25, 2022, 3:29pm UTC](https://es.discourse.group/t/code-point-string/1351/2 "2022-05-25T15:29:07Z")

</div>

I'm curious about how this would work in JSON. You can't add a raw string at the beginning of a .json file and have it still be valid. If there was a value of `"\u{10ffff}"` imported to js or otherwise, would it be `U+10FFFF` or the current value of `"\u{10ffff}"`?

---

<div class="post-metadata">

**Author:** ![Klaider](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/klaider/32/1447_2.png) [@Klaider](https://es.discourse.group/u/Klaider)\
**Post date:** [May 25, 2022, 3:31pm UTC](https://es.discourse.group/t/code-point-string/1351/3 "2022-05-25T15:31:35Z")

</div>

JSON doesn't manipulate, so that idea doesn't apply to JSON nor does it apply to literal string. When JSON is parsed, `\u{10ffff}` correctly turns into high-surrogate-n-low-surrogate format, or otherwise in UTF-8 format.

Ah, and JSON doesn't even support this `{}` escape, but it doesn't determine the encoding of the JSON string.

> If there was a value of `"\u{10ffff}"` imported to js or otherwise, would it be `U+10FFFF` or the current value of `"\u{10ffff}"` ?

The idea is that `'\u{10ffff}'` will be a single-character string with U+10FFFF. But, well, the encoding depends on the JavaScript engine.

* * *

So I did look at [V8](https://github.com/danbev/learning-v8/blob/d59d8ba5d4b4dd2c56b1be63507e020d4ab789e8/notes/string.md), seems like it supports both one-byte and two-byte encodings. It'd only require four-byte encoding for this feature to work.

---

<div class="post-metadata">

**Author:** ![ljharb](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/ljharb/32/8_2.png) [@ljharb](https://es.discourse.group/u/ljharb)\
**Post date:** [May 25, 2022, 4:18pm UTC](https://es.discourse.group/t/code-point-string/1351/4 "2022-05-25T16:18:44Z")

</div>

JS strings already have `.codePointAt` which I believe works closer to the way you expect?

---

<div class="post-metadata">

**Author:** ![Klaider](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/klaider/32/1447_2.png) [@Klaider](https://es.discourse.group/u/Klaider)\
**Post date:** [May 25, 2022, 4:22pm UTC](https://es.discourse.group/t/code-point-string/1351/5 "2022-05-25T16:22:00Z")

</div>

The problem of `.codePointAt` is that it receives an index based on code units... so the following would fail:

```javascript
// index '1' means 'second' character
'\u{10ffff}a'.codePointAt(1); // instead of U+61 ("a"), gives U+DFFF

```

---

<div class="post-metadata">

**Author:** ![ljharb](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/ljharb/32/8_2.png) [@ljharb](https://es.discourse.group/u/ljharb)\
**Post date:** [May 25, 2022, 4:33pm UTC](https://es.discourse.group/t/code-point-string/1351/6 "2022-05-25T16:33:17Z")

</div>

`Array.from(str)[1]` then.

---

<div class="post-metadata">

**Author:** ![Klaider](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/klaider/32/1447_2.png) [@Klaider](https://es.discourse.group/u/Klaider)\
**Post date:** [May 25, 2022, 4:36pm UTC](https://es.discourse.group/t/code-point-string/1351/7 "2022-05-25T16:36:49Z")

</div>

Fine, and then `String.fromCodePoint(...array)` back. But this would be inefficient if the string is long, especially when working with parsing.

---

<div class="post-metadata">

**Author:** ![theScottyJam](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/thescottyjam/32/616_2.png) [@theScottyJam](https://es.discourse.group/u/theScottyJam)\
**Post date:** [May 25, 2022, 11:45pm UTC](https://es.discourse.group/t/code-point-string/1351/8 "2022-05-25T23:45:14Z")

</div>

I believe the [iterator-helpers](https://github.com/tc39/proposal-iterator-helpers) proposal would help with the performance issues. If you're dealing with larger strings, you can write this:

```javascript
const [char] = Iterator.from(str).drop(index - 1);

```

But, it is unfortunately getting a bit more complicated now. Though, the solution to this complexity could be more iterator helpers.
