# Code point iterators for strings?

**URL:** <https://es.discourse.group/t/code-point-iterators-for-strings/206>\
**Category:** 💡 Ideas\
**Created:** [January 24, 2020, 8:10am UTC](https://es.discourse.group/t/code-point-iterators-for-strings/206 "2020-01-24T08:10:55Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![claudiameadows](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/claudiameadows/32/126_2.png) [@claudiameadows](https://es.discourse.group/u/claudiameadows)\
**Post date:** [January 24, 2020, 8:10am UTC](https://es.discourse.group/t/code-point-iterators-for-strings/206/1 "2020-01-24T08:10:55Z")

</div>

When processing strings in performance-sensitive code, you inevitably end up writing something like one of these two:

```javascript
// For state machine-based loops and similar
for (let i = 0; i < string.length; i++) {
    let code = string.charCodeAt(i)
    // ...
}

// For recursive-descent parsers
function next() {
    return index === string.length
        ? -1
        : string.charCodeAt(index++)
}

```

If you need full Unicode support, that invariably gets slightly more complicated:

```javascript
// For state machine-based loops and similar
for (let i = 0; i < string.length;) {
    let code = string.codePointAt(i++)
    if (code >= 0x10000) i++
    // ...
}

// For recursive-descent parsers
function next() {
    if (index === string.length) return -1
    let code = string.codePointAt(index++)
    if (code >= 0x10000) index++
    return code
}

```

Problem is, most engines store their strings not as simple byte sequences, so ‘String.prototype.charCodeAt`and`String.prototype.codePointAt`are *not* constant time. So could two new iterator methods be added to`String.prototype`?

- `String.prototype.codePoints()` - Iterate all code points in this string
- `String.prototype.charCodes()` - Iterate all character codes in this string

Each of these two would be relatively straightforward to define in JS:

```javascript
String.prototype.charCodes = function *() {
    let s = "" + this
    for (let i = 0; i < s.length; i++) {
        yield s.charCodeAt(i)
    }
}

String.prototype.codePoints = function *() {
    let s = "" + this
    for (let i = 0; i < s.length; i++) {
        let code = s.codePointAt(i)
        if (code > 0x10000) i++
        yield code
    }
}

```

An implementation might choose to optimize these to be a fully linear traversal, though, and they could optimize the iterator similarly to how they do with the default iterator.

Note that the code point iterator is likely to see greater use as most parsers use just that, but smaller use cases might not care about surrogates, and so they could skip the overhead.

---

<div class="post-metadata">

**Author:** ![claudepache](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/claudepache/32/236_2.png) [@claudepache](https://es.discourse.group/u/claudepache)\
**Post date:** [January 24, 2020, 8:38am UTC](https://es.discourse.group/t/code-point-iterators-for-strings/206/2 "2020-01-24T08:38:55Z")

</div>

Note that ECMAScript has already [`String.prototype[Symbol.iterator]()`](https://tc39.es/ecma262/#sec-string.prototype-@@iterator) that iterates over code points (although I think it was a blunder~~: it should have been named `String.prototype.codePoints()`~~).

EDIT: I realise after having written my comment that `String.prototype[Symbol.iterator]()` does not yields the information in the format you want, namely as integer rather than as 1-or-2-byte-string. (But, I’m still thinking it was a blunder.)

---

<div class="post-metadata">

**Author:** ![jridgewell](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/jridgewell/32/18_2.png) [@jridgewell](https://es.discourse.group/u/jridgewell)\
**Post date:** [January 24, 2020, 8:39am UTC](https://es.discourse.group/t/code-point-iterators-for-strings/206/3 "2020-01-24T08:39:46Z")

</div>

> [@claudiameadows](#):
>
> String.prototype.codePoints

> **[GitHub - tc39/proposal-string-prototype-codepoints:...](https://github.com/tc39/proposal-string-prototype-codepoints)**
>
> String.prototype.codePoints proposal for ECMAScript (stage 1) - GitHub - tc39/proposal-string-prototype-codepoints: String.prototype.codePoints proposal for ECMAScript (stage 1)

---

<div class="post-metadata">

**Author:** ![claudiameadows](https://yyz2.discourse-cdn.com/free1/user_avatar/es.discourse.group/claudiameadows/32/126_2.png) [@claudiameadows](https://es.discourse.group/u/claudiameadows)\
**Post date:** [January 24, 2020, 5:27pm UTC](https://es.discourse.group/t/code-point-iterators-for-strings/206/4 "2020-01-24T17:27:45Z")

</div>

I had not seen that before, so very interesting.

Is there any corresponding proposal for UCS-2 character codes?
