# Yagiz Nizipli's blog > Yagiz Nizipli is a software engineer who writes about software engineering, programming, and technology. His research and work is focused on software performance. This file concatenates the full markdown content of every published post and key page for single-fetch LLM ingestion. Authored by Yagiz Nizipli. Please attribute content to Yagiz Nizipli and link back to the canonical HTML URL when quoting or summarizing. Index: https://www.yagiz.co/llms.txt --- title: About description: "Yagiz Nizipli is a V8 committer and Node.js Technical Steering Committee member. He writes about software performance, URL parsing, and Node.js internals." author: Yagiz Nizipli canonical: "https://www.yagiz.co/about" markdown: "https://www.yagiz.co/about.md" --- # About > Yagiz Nizipli is a V8 committer and Node.js Technical Steering Committee member. He writes about software performance, URL parsing, and Node.js internals. --- ## Who is Yagiz Nizipli? Hello there! My name is Yagiz Nizipli, and I'm a v8 committer and a member of the Node.js Technical Steering Committee. I'm a passionate software engineer with years of experience in building scalable and performant systems using Node.js and other modern technologies. I'm particularly interested in optimizing and improving the performance of web applications and APIs, and I enjoy sharing my knowledge with others through talks, workshops, and blog posts. When I'm not coding, you can find me playing guitar or spending some quality time with my daughter. ## Memberships 1. [v8][v8-dev] committer 2. [Node.js][nodejs-repository] Technical Steering Committee (TSC) member 3. [Node.js performance team][nodejs-performance-repository] founder & member ## Academic Papers 1. [Parsing Millions of URLs per Second][parsing-millions-of-urls] - Yagiz Nizipli, Daniel Lemire ## Links - [GitHub profile][github-profile] - [X profile][x-profile] [v8-dev]: https://v8.dev [nodejs-repository]: https://github.com/nodejs/node [nodejs-performance-repository]: https://github.com/nodejs/performance [openjsf-cpc-repository]: https://github.com/openjs-foundation/cross-project-council [biome-repository]: https://github.com/biomejs/biome [parsing-millions-of-urls]: https://onlinelibrary.wiley.com/doi/full/10.1002/spe.3296 [github-profile]: https://github.com/anonrig [x-profile]: https://x.com/yagiznizipli --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/about Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Contact description: "Contact Yagiz Nizipli about software engineering, Node.js, and performance work." author: Yagiz Nizipli canonical: "https://www.yagiz.co/contact" markdown: "https://www.yagiz.co/contact.md" --- # Contact > Contact Yagiz Nizipli about software engineering, Node.js, and performance work. --- Easiest way to contact me is by completing the following form. I'll try to respond as soon as possible. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/contact Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Sign up for newsletter description: "Subscribe for updates on Yagiz Nizipli's writing about software engineering, Node.js, and performance." author: Yagiz Nizipli canonical: "https://www.yagiz.co/newsletter" markdown: "https://www.yagiz.co/newsletter.md" --- # Sign up for newsletter > Subscribe for updates on Yagiz Nizipli's writing about software engineering, Node.js, and performance. --- Please sign up for the newsletter to get notified about the latest articles, blog posts, and press releases featuring me and my projects. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/newsletter Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Press description: "Latest articles, blog posts and press releases featuring Yagiz Nizipli and his projects" author: Yagiz Nizipli canonical: "https://www.yagiz.co/press" markdown: "https://www.yagiz.co/press.md" --- # Press > Latest articles, blog posts and press releases featuring Yagiz Nizipli and his projects --- Latest articles, and press releases featuring me and my projects ## Articles 1. [Node.js Community Debate Intensifies Over Enabling Corepack][nodejs-community-debate-intensifies], Socket Security, 07/02/2024 2. [How to resolve unexpected request issues in a Node.js 20 upgrade][how-to-resolve-unexpected-request-issues], Tubi Engineering, 28/09/2023 3. [Bun: lessons from disrupting a tech ecosystem][bun-lessons-from-disrupting-tech], Pragmatic Engineer, 15/09/2023 4. [Node.js vs. Deno vs. Bun: JavaScript runtime comparison][nodejs-vs-deno-vs-bun-runtime-comparison], Snyk, 05/09/2023 5. [State of Node.js Performance 2023][state-of-nodejs-performance-2023], Rafael Gonzaga, 16/05/2023 6. [URL parser speed with Yagiz Nizipli][url-parser-speed], Daniel Steinberg, 04/02/2023 7. [Nominating @anonrig as a collaborator][nominating-anonrig-to-nodejs], Node.js, 06/10/2022 8. [Node Weekly, Issue #453][node-weekly-453], 09/09/2022 9. [Node Weekly, Issue #437][node-weekly-437], 15/05/2022 10. [JavaScript to Rust][javascript-to-rust], Node Weekly, 05/03/2022 11. [In Our Inspired Minds Today: Yagiz Nizipli][in-our-inspired-minds], Terminal Ventures, 31/05/2021 12. [Socketkit: Subscription & review analytics for iOS][socketkit-producthunt], Producthunt, 28/05/2021 13. [Mültecilerin sağlığı HERA’ya emanet][eczacinin-sesi], ECZACININ SESİ, 01/05/2021 14. [Socketkit: Kişisel verilerin güvenliğini sağlayarak veri toplayan analiz platformu][socketkit-webrazzi], Webrazzi, 29/04/2021 15. [Socketkit: Kişisel gizliliği koruyarak kullanıcılarınızdan veri toplamanızı sağlayan SaaS][socketkit-egirisim], E-Girişim, 28/04/2021 16. [Mültecilerin sağlığı HERA'ya emanet][hera-milliyet], Milliyet Gazetesi, 12/04/2021 17. [Mültecilerin sağlığı HERA’ya emanet][hera-msn], MSN, 11/04/2021 18. [Slack Manager: 'Agile' toplantılarınızın sanal yöneticisi][slack-manager-webrazzi], Webrazzi, 26/04/2016 19. [Google GDG DevFest Istanbul Etkinligi][gdg-wordpress], Wordpress.com, 06/12/2014 20. [Yağız Nizipli to speak at DevFest Istanbul][sabanci-devfest-en], Sabanci University, 28/11/2014 21. [Yağız Nizipli DevFest İstanbul'da Konuşacak!][sabanci-devfest-tr], Sabanci University, 28/11/2014 22. [Ulaşmak ve ulaşılır olmak][sabanci-varmigelen], Sabanci University, 28/11/2012 ## Printed Media 1. [**Kişisel Gizliliği Korunma Platformu**][printed-media-socketkit-milliyet], Milliyet, 14/04/2021 2. [**Mültecilerin sağlığı HERA'ya emanet**][printed-media-hera-milliyet], Milliyet, 14/04/2021 ## Presentations & Podcasts 1. [How Sentry debugs with Sentry][how-sentry-debugs-with-sentry], Youtube, 10/07/2024 2. [Optimizing URL parsing in Node.js][optimizing-url-parsing-in-nodejs], Podrocket, 04/10/2024 3. [Understanding package resolution in Node.js][understanding-package-resolution-in-nodejs], Node Congress, 04/04/2024 4. [Parsing millions of URLs per second][node-congress-parsing-millions-of-urls], Node Congress, 04/04/2024 5. [What is Node.js with Yagiz][what-is-nodejs-with-yagiz], This is Tech Talks, 19/02/2024 6. [JS Perf Wins & New Node.js Features with Yagiz Nizipli][talks-js-perf-wins], Syntax.fm, 10/01/2024 7. [From Rust to Parenthood: Yagiz Nizipli's Journey through Node.js, URL Parsing, and Personal Growth][presentation-from-rust-to-parenthood], Youtube, 18/12/2023 8. [Yagiz Nizipli - Node.js Performance][presentation-devtools-nodejs-performance-talk-2023], DevTools, 04/12/2023 9. [Ada: Parsing millions of URLs per second][presentation-ada-parsing-millions-of-urls], NodeConf, 07/11/2023 10. [Node.js Performance WG Session][presentation-nodejs-performance-2023], OpenJS Collaborator Summit, 15/05/2023 11. [Implementing a performant URL parser from scratch][presentation-conf42], Conf42, 17/11/2022 12. [Road to a fast URL parser in Node.js][presentation-nodeconf], NodeConf, 10/05/2022 13. [Inter Microservice Communication Patterns Where Latency Matters][presentation-feistanbul], Frontend Istanbul, 06/08/2021 14. [Open JOGL x HERA (Health Recording App)][presentation-openjogl], Youtube, 22/11/2019 15. [Developing scalable microservices][presentation-developing-scalable], Youtube, 04/02/2017 16. [Building microservices][presentation-building-ms], Youtube, 22/12/2016 17. [Building Scalable Microservices][presentation-kommunity], Kommunity, 01/12/2016 18. [Web Development: Making it the right way][presentation-web-dev], Youtube, 03/02/2015 ## Contributions 1. [Notable contributions on Node.js][nodejs-notable-contributions] [nodejs-community-debate-intensifies]: https://socket.dev/blog/node-community-debates-enabling-corepack-unbundling-npm [how-to-resolve-unexpected-request-issues]: https://code.tubitv.com/how-to-resolve-unexpected-request-issues-in-a-node-js-20-upgrade-31bf411058ba [bun-lessons-from-disrupting-tech]: https://blog.pragmaticengineer.com/bun-lessons-from-disrupting/ [nodejs-vs-deno-vs-bun-runtime-comparison]: https://snyk.io/blog/javascript-runtime-compare-node-deno-bun/ [state-of-nodejs-performance-2023]: https://blog.rafaelgss.dev/state-of-nodejs-performance-2023 [url-parser-speed]: https://lists.haxx.se/pipermail/daniel/2023-February/000002.html [nominating-anonrig-to-nodejs]: https://github.com/nodejs/node/issues/44906 [node-weekly-453]: https://nodeweekly.com/issues/453 [node-weekly-437]: https://nodeweekly.com/issues/437 [javascript-to-rust]: https://nodeweekly.com/issues/427 [in-our-inspired-minds]: https://us7.campaign-archive.com/?u=4cdadd41b1ec61abfaaa39a1c&id=6f6873e1c6&ref=yagiz.co [socketkit-producthunt]: https://www.producthunt.com/posts/subscription-review-analytics-for-ios [eczacinin-sesi]: https://eczacininsesi.com/haber-detay.php?id=13867&ref=yagiz.co [socketkit-webrazzi]: https://webrazzi.com/2021/04/29/socketkit-kisisel-verilerin-guvenligini-saglayarak-veri-toplayan-analiz-platformu [socketkit-egirisim]: https://egirisim.com/2021/04/29/socketkit-kisisel-gizliligi-koruyarak-kullanicilarinizdan-veri-toplamanizi-saglayan-saas/ [hera-milliyet]: https://www.milliyet.com.tr/gundem/multecilerin-sagligi-heraya-emanet-6479369 [hera-msn]: https://www.msn.com/tr-tr/haber/gundem/m%C3%BCltecilerin-sa%C4%9Fl%C4%B1%C4%9F%C4%B1-heraya-emanet/ar-BB1fygBI [slack-manager-webrazzi]: https://webrazzi.com/2016/04/26/slack-manager-agile-toplantilarinizin-sanal-yoneticisi-yerli-github [gdg-wordpress]: https://emrullahsaku.wordpress.com/2014/12/06/google-gdg-deffest-istanbul-etkinligi/ [sabanci-devfest-en]: https://gazetesu.sabanciuniv.edu/en/our-student-yagiz-nizipli-speak-devfest-istanbul [sabanci-devfest-tr]: https://gazetesu.sabanciuniv.edu/tr/ogrencimiz-yagiz-nizipli-devfest-istanbulda-konusacak [sabanci-varmigelen]: https://gazetesu.sabanciuniv.edu/tr/ulasmak-ve-ulasilir-olmak [how-sentry-debugs-with-sentry]: https://www.youtube.com/watch?v=yyaBjTCaJhA [optimizing-url-parsing-in-nodejs]: https://podrocket.logrocket.com/optimizing-url-parsing-in-node-yagiz-nizipli [understanding-package-resolution-in-nodejs]: https://portal.gitnation.org/contents/understanding-package-resolution-in-nodejs [node-congress-parsing-millions-of-urls]: https://portal.gitnation.org/contents/parsing-millions-of-urls-per-second [what-is-nodejs-with-yagiz]: https://www.youtube.com/watch?v=c6DRXZk9S9I [talks-js-perf-wins]: https://syntax.fm/show/716/js-perf-wins-and-new-node-js-features-with-yagiz-nizipli [presentation-from-rust-to-parenthood]: https://www.youtube.com/watch?v=1ex8dZ7i8_M [presentation-devtools-nodejs-performance-talk-2023]: https://www.youtube.com/watch?v=waEKI4V70CU [presentation-ada-parsing-millions-of-urls]: https://www.youtube.com/watch?v=tQ-6OWRDsZg [presentation-nodejs-performance-2023]: https://www.youtube.com/watch?v=D7gd2pLo_QY [presentation-conf42]: https://www.conf42.com/JavaScript_2022_Yagiz_Nizipli_performant_url_parser [presentation-nodeconf]: https://www.youtube.com/watch?v=hT25FOx5kNQ&ref=yagiz.co [presentation-feistanbul]: https://kommunity.com/frontend-istanbul/events/inter-microservice-communication-patterns-where-latency-matters-1de2dec0 [presentation-openjogl]: https://www.youtube.com/watch?v=4KKbh6xk3EU&ref=yagiz.co [presentation-developing-scalable]: https://www.youtube.com/watch?v=1XLxHg6XRjw&ref=yagiz.co [presentation-building-ms]: https://www.youtube.com/watch?v=HIKWzjSwJhI&ref=yagiz.co [presentation-kommunity]: https://kommunity.com/istanbulcoders/events/236332031 [presentation-web-dev]: https://www.youtube.com/watch?v=0qbEZ8PBNm0&ref=yagiz.co [nodejs-notable-contributions]: https://github.com/nodejs/node/pulls?q=is%3Apr+author%3Aanonrig+label%3Anotable-change+is%3Aclosed [printed-media-socketkit-milliyet]: https://www.yagiz.co/press/socketkit-milliyet-2021.jpeg [printed-media-hera-milliyet]: https://www.yagiz.co/press/hera-milliyet-2021.jpeg --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/press Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: URL parsing description: "A series on WHATWG URL parsing, Ada, Node.js, and the SIMD and serialization work behind them." author: Yagiz Nizipli canonical: "https://www.yagiz.co/url-parsing" markdown: "https://www.yagiz.co/url-parsing.md" --- # URL parsing > A series on WHATWG URL parsing, Ada, Node.js, and the SIMD and serialization work behind them. --- - [Implementing Node.js URL parser in WebAssembly with Rust](https://www.yagiz.co/implementing-node-js-url-parser-in-webassembly-with-rust.md) — 2022-02-28 - [Working with 1000+ tests on a stable C++ library](https://www.yagiz.co/working-with-1000-tests.md) — 2023-03-18 - [Announcing Ada URL parser v2.0](https://www.yagiz.co/announcing-ada-url-parser-v2-0.md) — 2023-03-30 - [Reducing the cost of string serialization in Node.js core](https://www.yagiz.co/reducing-the-cost-of-string-serialization-in-nodejs-core.md) — 2023-04-25 - [URL specification and browser implementation differences](https://www.yagiz.co/url-parsing-and-browser-differences.md) — 2023-05-25 - [The story of WHATWG URL Specification for toddlers](https://www.yagiz.co/whatwg-url-specification-for-toddlers.md) — 2023-06-22 - [Release of Ada v3.0 with URLPattern](https://www.yagiz.co/release-of-ada-v3.md) — 2025-01-30 - [State of URL parsing performance in 2025](https://www.yagiz.co/state-of-url-parsing-2025.md) — 2025-12-09 - [Announcing Ada v4: Validating 35.6M URLs per second](https://www.yagiz.co/release-of-ada-v4.md) — 2026-07-27 - [192.168.1.1 is too short for SIMD](https://www.yagiz.co/simd-is-the-wrong-way-to-parse-ipv4.md) — 2026-08-24 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/url-parsing Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: 192.168.1.1 is too short for SIMD description: "An IPv4 host is 7 to 16 bytes, so it never fills a SIMD register. I unroll the scalar parse, try a SWAR fold that loses, put SSE2 in front of it, and only then reach for AVX-512." date: 2026-08-24 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/simd-is-the-wrong-way-to-parse-ipv4" markdown: "https://www.yagiz.co/simd-is-the-wrong-way-to-parse-ipv4.md" --- # 192.168.1.1 is too short for SIMD > An IPv4 host is 7 to 16 bytes, so it never fills a SIMD register. I unroll the scalar parse, try a SWAR fold that loses, put SSE2 in front of it, and only then reach for AVX-512. *Published: 2026-08-24 · Tag: performance* --- Suppose you want to parse `192.168.1.1` as a URL host. In [Ada][ada] that string is not a domain name. It is an IPv4 address, and the [WHATWG URL Standard][whatwg-ipv4] has a whole parser for it. A reasonable function might look as follows. ```cpp title="The obvious IPv4 parser" uint64_t parse_ipv4(std::string_view s) { const char* p = s.data(); const char* end = p + s.size(); uint32_t addr = 0; for (int i = 0; i < 4; ++i) { if (p == end || *p < '0' || *p > '9') { return fail; } uint32_t val = uint32_t(*p++ - '0'); while (p < end && *p >= '0' && *p <= '9') { val = val * 10 + uint32_t(*p++ - '0'); if (val > 255) { return fail; } } addr = (addr << 8) | val; if (i < 3) { if (p == end || *p != '.') { return fail; } ++p; } } return (p == end) ? addr : fail; } ``` Four octets, a dot between them. Out of range or out of place returns `fail`. Consume the whole string and you get a packed `uint32_t` in host order, first octet in the high byte. `192.168.1.1` becomes `0xC0A80101`. This version also treats `01.2.3.4` as `1.2.3.4`. WHATWG does not. A leading zero starts an octal number. I will come back to that. `0.0.0.0` is 7 bytes. `255.255.255.255` is 15. WHATWG also allows one trailing dot (`1.2.3.4.`). Ada's fast path accepts that trailing dot, not a second one. I generated 10,000 random addresses with each octet uniform in 0 to 255. Average length 13.3 bytes. Shortest 9, longest 15. A 128-bit register holds 16 bytes, so you are always leaving lanes unused. In my [last post][branches] the input was a megabyte and SIMD was the right tool. Here I am not sure the setup is cheaper than four short decimal conversions. [Daniel Lemire][lemire] has a [post][lemire-ipv4] on doing this with SIMD, and a [follow-up][lemire-csharp] that revisits it with AVX-512. Wojciech Muła has an [article][mula] on the same problem. I care about the 12-byte case, and about still having a general parser for the weird forms. ## Hex, octal, and `127.1` WHATWG IPv4 is not just `ddd.ddd.ddd.ddd`. Browsers also accept `127.1`, `127.0.1`, `0x7f.0.0.1`, `0177.0.0.1`, and `3232235777`. Ada has to accept them too. That parser has to detect `0x`, reject `08` as a bad octal, and shift the last part by 8, 16, or 24 bits depending on how many dots it saw. glibc `inet_pton` is a different grammar. It wants a nul-terminated C string, writes an `in_addr`, and does not implement WHATWG. I use it as a baseline. The common case is still `192.168.0.1` or `12.121.244.111`. Four decimal numbers, no leading zeros, optional trailing dot. Everything else can fall through. ```cpp title="Fast path, then the general parser" if (uint64_t ip = try_parse_ipv4_fast(s); ip < (1ull << 32)) { return ip; } return parse_ipv4_general(s); ``` Daniel uses the same split in the [C# post][lemire-csharp]. One rule the fast path must get right: a leading zero is not decimal. `01.2.3.4` is octal in WHATWG. If the first digit of a group is `0` and a second digit follows, we return `fail` and let the general parser decide. `0.0.0.0` is fine. The zero is the whole group. ## Unroll the four octets The obvious parser has a `while` inside the `for`. Each octet may be 1, 2, or 3 digits, and the inner loop branches on a value the CPU has just loaded. We know there are exactly four groups, and each group is at most three digits. We can write that down. ```cpp title="Unrolled pure-decimal IPv4" uint64_t parse_ipv4_unrolled(const char* p, const char* end) { uint32_t addr = 0; for (int i = 0; i < 4; ++i) { if (p == end || *p < '0' || *p > '9') { return fail; } uint32_t val = uint32_t(*p++ - '0'); if (p < end && *p >= '0' && *p <= '9') { if (val == 0) { return fail; // leading zero: not our problem } val = val * 10 + uint32_t(*p++ - '0'); if (p < end && *p >= '0' && *p <= '9') { val = val * 10 + uint32_t(*p++ - '0'); if (val > 255) { return fail; } } } addr = (addr << 8) | val; if (i < 3) { if (p == end || *p != '.') { return fail; } ++p; } } if (p == end) { return addr; } if (p + 1 == end && *p == '.') { return addr; // WHATWG trailing dot } return fail; } ``` Let's walk through `192.168.1.1`: - First group: `'1'`, then `'9'`, then `'2'`. `val` becomes 192. Next byte is `'.'`. Shift 192 into the address. - Second group: 168. Same shape. - Third group: `'1'`. The next byte is `'.'`, so we stop at one digit. No leading-zero check, because there is no second digit. - Fourth group: `'1'`. `p == end`. Done. `10.0.0.1` is the same, with more one-digit groups. `01.2.3.4` hits `val == 0` on the second digit of the first group and returns `fail`. The general parser then reads it as octal. There is no inner `while`. Each extra digit is a predicted forward branch. This is the portable path in Ada, `parse_ipv4_decimal_scalar` in [`checkers-inl.h`][checkers], and what we ship when the machine does not have AVX-512. Can we do better? ## A SWAR fold that loses In the last post, a 256-byte table lost to a wraparound compare. I tried the same kind of trick here. Once you know a group is three digits, you can load them as a `uint32_t`, subtract `'0'` from each byte, reverse the digits so the ones place sits in the low byte, and fold with the weights `1, 10, 100`. Those weights are the constant `0x0000640a01`. ```cpp title="Ones-first fold of one octet" // "192" reversed is bytes 2, 9, 1, 0. // 0x640a01 is the weights {1, 10, 100, 0}. uint32_t fold(uint32_t ones_first) { return (ones_first & 0xff) + 10u * ((ones_first >> 8) & 0xff) + 100u * ((ones_first >> 16) & 0xff); } ``` `255` becomes the bytes `5, 5, 2, 0`. That dword is `0x00020505`. `256` becomes `6, 5, 2, 0`, which is `0x00020506`. The unsigned compare `dword > 0x00020505` is the range check. We will use that layout again in a minute. I wired this fold into the unrolled parser. It was twice as slow as multiplying by ten as you go. The load, the reverse, and the three-term add are extra work on a group that is already two or three bytes long. SWAR wants eight bytes in a register. An IPv4 octet is not eight bytes. Keep the incremental `* 10`. The ones-first dword is still a good way to compare against 255. It is a bad way to convert a single octet in scalar code. ## Why SSE2 pre-validation does not help The first SIMD version I would write is not a parser. It is a pre-check. Load 16 bytes, ask whether every live byte is a digit or a dot, ask whether there are exactly three dots, then run the unrolled scalar parser. ```cpp title="SSE2 pre-validation, then the scalar parse" uint64_t parse_ipv4_sse2(const char* padded16, const char* exact, size_t len) { const __m128i v = _mm_loadu_si128(reinterpret_cast(padded16)); const __m128i digits = _mm_sub_epi8(v, _mm_set1_epi8('0')); const __m128i is_digit = _mm_cmpeq_epi8( _mm_min_epu8(digits, _mm_set1_epi8(9)), digits); const __m128i is_dot = _mm_cmpeq_epi8(v, _mm_set1_epi8('.')); const int live = int((1u << len) - 1u); const int ok = _mm_movemask_epi8(_mm_or_si128(is_digit, is_dot)); if ((ok & live) != live) { return fail; } const int dots = _mm_movemask_epi8(is_dot) & live; if (__builtin_popcount(unsigned(dots)) != 3) { return fail; } return parse_ipv4_unrolled(exact, exact + len); } ``` `padded16` is the annoying part. `_mm_loadu_si128` reads 16 bytes. A 12-byte host in the middle of a URL is not 16 bytes long. Load past the end of the allocation and you have undefined behavior. On the last page of a buffer it can fault. Daniel [ran into this in 2023][lemire-ipv4]. His `sse_inet_aton` always reads sixteen bytes. The input must be part of a larger string, or you must overallocate. [Wojciech's article][mula] has the same constraint. You can pad. You can copy the host into a 16-byte stack buffer. You can prove the URL buffer has slack. All of those cost something. The unrolled scalar path reads exactly `len` bytes. Even if you pad for free, the pre-check does not replace the parse. It adds a 16-byte load, three compares, a movemask, and a popcount in front of the function you were going to run anyway. On this machine that pair is a tie. [Wojciech's actual insight][mula] is better than a pre-check. There are only 81 ways to place three dots in a 16-byte window, plus a virtual dot at the end. You build a mask of the dot positions, map that mask through a table, and get a `pshufb` pattern that lays the digits out for a SIMD convert. His earlier version stops after locating the dots and converts each group in scalar code. I ran both. The production SSE convert I timed is the one in [simdzone][simdzone]. ## Masked loads and `vpcompressb` AVX-512 changes the setup cost. A masked load takes a 16-bit mask and reads only those bytes. If the host is 12 bytes long, you load 12 bytes. You do not read past the string. You do not pad. Daniel has a [post][masked] on this, and the [C# IPv4 parser][lemire-csharp] uses it to pull UTF-16 characters without over-reading. Once the bytes are in a register, VBMI2 gives you `vpcompressb`. That is the instruction that makes the 81-entry table unnecessary. You take the positions of the dots, compress them into a tight vector, and compute the digit shuffle from those positions. No lookup. The Ada kernel is the table-free AVX-512 path from [simdip][simdip] (`parse_ipv4_avx512vl_notab5`). It lives in [`try_parse_ipv4_fast`][checkers]. ```cpp title="Masked load of exactly len bytes" const uint32_t len_mask = _bzhi_u32(~0u, unsigned(len)); const __m128i v = _mm_mask_loadu_epi8( _mm_set1_epi8('.'), __mmask16(len_mask), data); ``` `_bzhi_u32` keeps the low `len` bits. The masked load fills lanes outside the mask with `'.'`. A padding dot cannot look like a digit, and the first padding dot sits at index `len`. It is the virtual fourth delimiter. Subtract `'0'` from every lane. Lanes that were digits become 0 to 9. A compare gives you a digit mask. Anything inside `len` that is neither a digit nor a dot is junk, and we fail. ```cpp title="Compress the dot positions" const __mmask16 delim = _mm_cmpeq_epi8_mask(v, _mm_set1_epi8('.')); const __m128i iota = _mm_setr_epi8( 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15); const __m128i c = _mm_maskz_compress_epi8(delim, iota); ``` Take `192.168.1.1`. The string is 11 bytes. The dots sit at 3, 7, and 9. The first padding dot sits at 11. After the compress, `c` begins `{3, 7, 9, 11, ...}`. From each delimiter, walk backward one, two, and three bytes. That is the ones, tens, and hundreds digits. Broadcast each delimiter into four lanes and subtract `{1, 2, 3, 4}`: ```cpp title="idx = max(qi + offr, prev)" const __m128i k_rep = _mm_setr_epi8( 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3); const __m128i qi = _mm_shuffle_epi8(c, k_rep); const __m128i prev = _mm_shuffle_epi8( _mm_alignr_epi8(c, _mm_set1_epi8(-1), 15), k_rep); const __m128i offr = _mm_setr_epi8( -1, -2, -3, -4, -1, -2, -3, -4, -1, -2, -3, -4, -1, -2, -3, -4); const __m128i idx = _mm_max_epi8(_mm_add_epi8(qi, offr), prev); ``` `qi` is the end of each group, repeated four times: `{3,3,3,3, 7,7,7,7, 9,9,9,9, 11,11,11,11}`. `prev` is the start of each group: `{-1,-1,-1,-1, 3,3,3,3, 7,7,7,7, 9,9,9,9}`. The `-1` is "before the first byte." `qi + offr` walks backward from the delimiter: ones, tens, hundreds, and an unused lane. `max(..., prev)` clamps that walk so a short group cannot steal a byte from the previous group. A missing tens or hundreds digit lands on `prev`, which is a dot, which we have already zeroed in the digit vector. `pshufb` of `-1` is also zero. For the first group of `192.168.1.1`: - ones: `max(3 - 1, -1) = 2`, the `'2'` - tens: `max(3 - 2, -1) = 1`, the `'9'` - hundreds: `max(3 - 3, -1) = 0`, the `'1'` For the third group, the single `'1'` at index 8: - ones: `max(9 - 1, 7) = 8` - tens: `max(9 - 2, 7) = 7`, the previous dot, so zero - hundreds: `max(9 - 3, 7) = 7`, zero again One `pshufb` gathers all four groups. Each group is stored as `ones | (tens << 8) | (hundreds << 16)`. `255` is the bytes `5, 5, 2, 0`, the dword `0x00020505`. `256` is `6, 5, 2, 0`. `192` is `2, 9, 1, 0`, the dword `0x00010902`. ```cpp title="Octet > 255 is one unsigned dword compare" const __m128i lim = _mm_setr_epi8( 5, 5, 2, 0, 5, 5, 2, 0, 5, 5, 2, 0, 5, 5, 2, 0); const __mmask8 over = _mm_cmpgt_epu32_mask(padded, lim); ``` Four groups, four dwords, one compare. We range-check the digits before converting them. The two can run in parallel. The convert itself is a 4-wide dot product of those bytes with `1, 10, 100, 0`. Same fold as `0x640a01`, four times, in one instruction. On Ice Lake and later that is `vpdpbusd`. Without VNNI it is `pmaddubsw` plus `pmaddwd`. Same answer. ```cpp title="1*ones + 10*tens + 100*hundreds, four groups at once" const __m128i wts = _mm_setr_epi8( 1, 10, 100, 0, 1, 10, 100, 0, 1, 10, 100, 0, 1, 10, 100, 0); const __m128i res = _mm_dpbusd_epi32(_mm_setzero_si128(), padded, wts); ``` A few more masks finish the validation. Exactly three live dots: `popcnt(dots) ^ 3` is zero only then, and padding dots do not count. Each group is 1 to 3 digits, which is a gap check on the compressed positions. No leading zeros: a `'0'` at the start of a group, with a digit after it, is a fail. `0.0.0.0` has no digit after those zeros, so it passes. If every mask is clean, we pack the four bytes and `bswap` into host order. If anything is off, including hex, octal, two-part addresses, or `256.0.0.1`, we return `fail` and the general parser runs. ## Numbers I ran these on an Intel Xeon with GCC 13.3 (`-O3 -march=native`). The machine is 2.4 GHz and has AVX-512BW, VL, VBMI2, and VNNI. I generated 10,000 random dotted-decimal addresses and parsed the set 2,000 times. Every input is valid. The 81-mask and Muła functions read 16 bytes. I zero-padded the corpus up front, so those numbers do not include a `memcpy`. The scalar paths and the AVX-512 path read exactly the live bytes. | Approach | ns/addr | million addr/s | | --- | ---: | ---: | | `inet_pton` | 30 | 34 | | Unrolled + SWAR fold | 35 | 29 | | Naive loop | 17 | 59 | | Unrolled scalar | 16 | 62 | | SSE2 pre-check + unrolled | 16 | 61 | | Muła SSE (locate + scalar) | 15 | 66 | | AVX-512 VBMI2 | 4.5 | 220 | | simdzone 81-mask SSE | 3.1 | 320 | `inet_pton` is doing more work than we need: a nul-terminated string, an `in_addr` write, and a grammar that is not WHATWG. The naive loop is already twice `inet_pton`. Unrolling the digit counts shaves a little more. The SWAR fold loses to that unrolled loop by about 2x. SSE2 in front of the unrolled loop is a tie. Sometimes it loses by a nanosecond, sometimes it matches. I would not take a 16-byte over-read for that. Muła's locate-then-scalar SSE is a small win over unrolled, when the pad is free. It still converts each group in scalar code. The convert, not the locate, is the expensive part. The [simdzone][simdzone] 81-mask parser is the real SIMD convert. On this machine it is the fastest happy-path function I ran, faster than the AVX-512 kernel. Not a surprise: the input is already 16-byte padded, every address is valid, the table lookup is branchless, and there is no masked-load tax. Do not line that 320 up against [Daniel's 2023 Ice Lake number][lemire-ipv4]. Different machine, different compiler, different padding contract. AVX-512 is four times the unrolled scalar path, and it does not need the pad. That is why Ada ships it. A URL host is a slice of a larger string. I do not want to promise 16 readable bytes at the end of a buffer. This sits under `parse_host`. Faster IPv4 helps when the host is IPv4. It does nothing for `https://example.com/`. A host that starts with a letter never reaches this function. I also timed the miss path. On a corpus of only the unusual WHATWG forms (`0x7f.0.0.1`, `0177.0.0.1`, `127.1`, `01.2.3.4`, a 32-bit decimal, and so on): | Approach | ns/addr | | --- | ---: | | Unrolled scalar (rejects) | 2.6 | | General parser only | 10 | | AVX-512, then general | 11 | The fast path is cheap to fail. Running it before the general parser costs about a nanosecond on a corpus where every address misses. On a mixed set, 19 dotted-decimal addresses for every unusual one, AVX-512-then-general stays at 4.6 ns. The general parser alone is 22 ns, because it is now doing the common case the hard way. ## Which one should you use? It depends on the machine and on how much IPv4 you actually see. 1. **Unrolled scalar** as the default. Four groups, at most three digits each, leading zeros rejected. No over-read. No table. This is what Ada uses when AVX-512 VBMI2 is not available. 2. **Do not put SSE2 or NEON in front of that loop** to "pre-validate" a 7-16 byte host. You still have to parse, and you have now promised to read 16 bytes. On this machine that pair is a tie. 3. **Do not SWAR-fold a three-digit octet.** The ones-first dword is a compare trick, not a convert trick. Incremental `* 10` won. 4. **The 81-mask SSE tables** if IPv4 is the whole program and you can overallocate. [Daniel][lemire-ipv4] and [Wojciech][mula] already wrote this. It was the fastest happy-path function on this Xeon. I would not rebuild it for a URL parser. 5. **AVX-512 VBMI2** when you have it and IPv4 shows up in the profile. Masked load, compress the dots, dword-compare against 255. Keep the general parser for everything the kernel refuses. 6. **Always keep the fallback.** Hex, octal, and `127.1` are not worth teaching to a SIMD kernel. A trailing dot is not one of those cases. The fast path already accepts it. Compilers will unroll a four-iteration loop. They will not invent a masked load, and they will not notice that `255` is the dword `5, 5, 2, 0`. So far every address was four decimal groups. Suppose you want a harder problem. ## A harder problem: IPv6 IPv6 is the 128-bit form, written as eight hex groups with colons: `2001:db8::1`. The `::` may appear once and means "fill the rest with zeros." You can also embed an IPv4 address at the end (`::ffff:192.168.1.1`). The string can be 2 bytes (`::`) or 39 (`2001:0db8:0000:0000:0000:0000:0000:0001`), or 45 if you write the embedded IPv4 in decimal. This is a different shape. The dots in IPv4 tell you where every group starts. The colons in IPv6 do not, because of `::`. A table of 81 masks does not exist here. You first have to decide whether the input is even plausible: how many colons, whether `::` appears twice, whether a dot means an embedded IPv4. Ada has an AVX-512 helper that only answers that question. One masked 512-bit load, a compare against `:` and `.`, a few popcounts. Impossible strings never reach the piece parser. Possible strings still go through the scalar walk. ```cpp title="IPv6 shape check, one masked 512-bit load" bool ipv6_structure_plausible(const char* data, size_t len) { if (len < 2 || len > 45) { return false; } const __mmask64 live = __mmask64((1ULL << len) - 1ULL); const __m512i input = _mm512_maskz_loadu_epi8(live, data); const __mmask64 is_colon = _mm512_mask_cmpeq_epi8_mask(live, input, _mm512_set1_epi8(':')); const __mmask64 is_dot = _mm512_mask_cmpeq_epi8_mask(live, input, _mm512_set1_epi8('.')); const int colons = int(_mm_popcnt_u64(uint64_t(is_colon))); if (colons > 8) { return false; } const uint64_t doubles = uint64_t(is_colon) & (uint64_t(is_colon) << 1); if (doubles != 0 && (doubles & (doubles - 1)) != 0) { return false; // more than one "::" } if (doubles == 0 && is_dot == 0 && colons != 7) { return false; } return true; } ``` `doubles` is the colon mask AND-shifted onto itself. Adjacent colons light a bit. If more than one bit is set, you have two `::` (or `:::`), and the string is impossible. A full form without `::` and without an embedded IPv4 must have exactly seven colons. Everything else, including `::1` and `::ffff:192.168.1.1`, is plausible and goes to the scalar piece parser. On a mix of valid addresses, two-`::` junk, an over-long form, and a domain name, the scalar scan is about 10 ns. The masked load is about 1.4 ns. Daniel has a [post][lemire-ipv6] on a full AVX-512 convert (Shreesh Adiga's kernel): one 512-bit load, expand on the colons, permute the hex. He gets about 70 million addresses per second on an Emerald Rapids Xeon, against `inet_pton`. That is a different job. In a URL parser the common host is still a domain name, and the cheapest parse is the one you do not run. In Ada, a host that starts with a letter never calls either IP parser. [ada]: https://github.com/ada-url/ada [checkers]: https://github.com/ada-url/ada/blob/main/include/ada/checkers-inl.h [whatwg-ipv4]: https://url.spec.whatwg.org/#concept-ipv4-parser [lemire]: https://lemire.me [lemire-ipv4]: https://lemire.me/blog/2023/06/08/parsing-ip-addresses-crazily-fast/ [lemire-csharp]: https://lemire.me/blog/2026/08/19/parsing-ip-addresses-in-c-at-crazy-speeds/ [lemire-ipv6]: https://lemire.me/blog/2026/05/23/parsing-ipv6-addresses-crazily-fast-with-avx-512/ [mula]: http://0x80.pl/notesen/2023-04-09-faster-parse-ipv4.html [masked]: https://lemire.me/blog/2022/11/08/modern-vector-programming-with-masked-loads-and-stores/ [simdip]: https://github.com/lemire/simdip [simdzone]: https://github.com/NLnetLabs/simdzone [branches]: https://www.yagiz.co/eliminating-branches-in-cpp-loops --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/simd-is-the-wrong-way-to-parse-ipv4 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Eliminating branches in C++ loops description: "A for loop that returns false from a branch and true at the end is easy to read, but the branch can be expensive. In this post I show how to remove it with bitwise operations, a 256-byte table, SWAR, and NEON, and then apply the same ideas to UTF-16 validation." date: 2026-08-22 tag: performance author: Yagiz Nizipli canonical: "https://www.yagiz.co/eliminating-branches-in-cpp-loops" markdown: "https://www.yagiz.co/eliminating-branches-in-cpp-loops.md" --- # Eliminating branches in C++ loops > A for loop that returns false from a branch and true at the end is easy to read, but the branch can be expensive. In this post I show how to remove it with bitwise operations, a 256-byte table, SWAR, and NEON, and then apply the same ideas to UTF-16 validation. *Published: 2026-08-22 · Tag: performance* --- Suppose you want to check whether a string is made entirely of ASCII lowercase letters. It is a common check in parsers. In [Ada][ada] we do this kind of classification constantly for URL characters: is this a hex digit, an unreserved character, a forbidden host code point? A reasonable function might look as follows. ```cpp title="The obvious validating loop" bool is_ascii_lowercase(std::string_view input) { for (unsigned char c : input) { if (c < 'a' || c > 'z') { return false; } } return true; } ``` If any byte falls outside `a-z`, we return `false`. If the loop finishes, we return `true`. Importantly, this function exits as soon as a bad character is found. If we expect that almost every input is valid, that early `return` can be expensive. The CPU guesses which side of the `if` will run. When it guesses wrong, you pay a pipeline flush. `||` and `&&` make it worse because they short-circuit: the second compare is itself a branch. [Daniel Lemire][lemire] has a [post][escaping] on a similar problem: checking whether a JSON string needs escaping. The structure is the same. A loop, a branch, `return true` at the end. The rest of this post follows the same ladder he uses there: scan the whole string, replace the compare with a table, then do eight or sixteen bytes at once. Always cast the byte to `unsigned char` (or `uint8_t`) before you classify it. A plain `char` may be signed, and a signed value above 127 can become a negative index or break the range check. ## Scan the whole string If we expect that no bad character will be found, we can always scan the whole input. That lets the compiler try other optimizations. In particular, it is more likely to autovectorize the loop: to compile it using SIMD instructions on its own. Daniel calls this version branchless, because it does not branch out of the loop. ```cpp title="Branchless accumulation" bool is_ascii_lowercase(std::string_view input) { bool ok = true; for (unsigned char c : input) { ok &= (c >= 'a') & (c <= 'z'); } return ok; } ``` `&` is not `&&`. Bitwise AND always evaluates both sides, so there is no short-circuit branch. The loop body is load, compare, compare, and, store. After the last byte we return the flag. I prefer the dual form when I am looking for problems rather than confirming that everything is valid. Accumulate errors with OR: ```cpp title="Branchless accumulation with OR" bool is_ascii_lowercase(std::string_view input) { unsigned errors = 0; for (unsigned char c : input) { errors |= static_cast(c < 'a'); errors |= static_cast(c > 'z'); } return errors == 0; } ``` On x86 those compares compile to `setcc`. That is a flag write, not a jump. The loop always runs to completion. That is what you want when the happy path is that the whole string is fine, and it is what the vectorizer wants to see. We still have two comparisons per byte. We can do better. ## One compare: the wraparound test A byte is an ASCII lowercase letter if and only if it sits in `['a', 'z']`. Subtract `'a'` and that statement becomes "the result fits in 0 to 25". ```cpp title="Range check that wraps into a single unsigned compare" static inline bool is_lower(unsigned char c) { return static_cast(c - 'a') <= 25; } ``` Let's walk through a few values: - `'a' - 'a'` is `0`, and `0 <= 25`. - `'z' - 'a'` is `25`, still in range. - `` '`' - 'a' `` wraps to `255`, and `255 <= 25` is false. - `'{' - 'a'` is `26`, just outside. - `'A' - 'a'` wraps well above 25, so uppercase is rejected. Bytes below `'a'` underflow modulo 256 and land in the high end of the `unsigned char` range, where they fail the same `<= 25` test as bytes above `'z'`. Two comparisons become one. ```cpp title="Validating a string with a wraparound predicate" bool is_ascii_lowercase(std::string_view input) { unsigned errors = 0; for (unsigned char c : input) { errors |= static_cast( static_cast(c - 'a') > 25); } return errors == 0; } ``` The body is now subtract, compare, or. There is no `if`, no `return` in the middle, and no `||`. For a single closed interval like `a-z`, this is usually as far as scalar code needs to go. The wraparound test only works for one interval. Hex digits, unreserved URL characters, and forbidden host code points are unions of ranges and punctuation. Bitwise arithmetic gets ugly there. A table does not. ## A 256-byte lookup table A simple way to classify a byte is to generate a 256-element array and look the value up. Daniel calls this [memoization][tables] (and not memorization). You will sometimes hear "a table of size 255". The last valid index of an 8-bit value is 255, but the length of the array is **256**. Byte values run from `0x00` through `0xFF` inclusive. A table of 255 entries leaves `0xFF` unmapped. Using C++17, you can have the compiler build the array at compile time from a lambda: ```cpp title="Build a 256-entry lowercase table at compile time" static constexpr std::array kIsLower = []() { std::array table{}; for (unsigned c = 'a'; c <= 'z'; ++c) { table[c] = 1; } return table; }(); bool is_ascii_lowercase(std::string_view input) { uint8_t ok = 1; for (unsigned char c : input) { ok &= kIsLower[c]; } return ok != 0; } ``` `table{}` zero-initializes every slot. The loop turns on only `'a'` through `'z'`. Everything else stays `0`. Each character is checked with a single load, plus an AND. This might compile down to a single lookup instruction. I am using lowercase here so the skeleton is easy to see. I would not ship a table for `a-z`. The wraparound test is already one subtract and one compare. A table replaces that with a load, and the load is slower. The table starts to win when the check is no longer a single interval. Hex is the example I actually use this for. Two ranges plus a gap: ```cpp title="Hex digits are a messy range and a clean table" static constexpr std::array kIsHex = []() { std::array table{}; for (unsigned c = '0'; c <= '9'; ++c) table[c] = 1; for (unsigned c = 'a'; c <= 'f'; ++c) table[c] = 1; for (unsigned c = 'A'; c <= 'F'; ++c) table[c] = 1; return table; }(); bool is_ascii_hex(std::string_view input) { uint8_t ok = 1; for (unsigned char c : input) { ok &= kIsHex[c]; } return ok != 0; } ``` The scalar version of that check is `(c >= '0' && c <= '9') || (c >= 'a' && c <= 'f') || (c >= 'A' && c <= 'F')`. Six comparisons and a tree of short-circuit branches, or one load. This is how Ada classifies URL characters. [Daniel's compile-time table post][tables] uses the same idea for forbidden host code points: `'\0'`, tab, space, `#`, `/`, `:`, and so on. Those are not a single interval. They are a pile of allowed and forbidden bytes. A 256-byte table per class is cheaper than explaining those rules to the branch predictor. Two caveats: - Index the table with `uint8_t` or `unsigned char`. A signed `char` of `0xFF` becomes `-1` and walks off the front of the array. - A table only wins if it stays in cache. 256 bytes is four cache lines. Rebuilding the table on every call is extra work, not a table. I will come back to this with numbers. The short version is that a table is the right tool for hex or forbidden host bytes, and the wrong tool for `a-z`. A range is already one subtract and one compare. A table turns that into a load, and loads have latency. Can we do better? ## SWAR when SIMD is not available If you do not have SIMD, or you do not want a runtime ISA dispatch, you can still process eight bytes at a time. The technique is called SWAR: SIMD within a register. [Lamport described it in 1975][swar]. The intuition is that modern computers have 64-bit registers. Processing eight consecutive bytes as eight distinct words is inefficient given how wide our registers are. The first step is to load eight characters into a `uint64_t`. In C++, you might do it this way: ```cpp title="Load eight bytes into a register" uint64_t word; std::memcpy(&word, chars, 8); ``` It looks maybe expensive, but most compilers will translate the `memcpy` into a single load when optimizations are on. We then treat that register as a vector of eight bytes and use ordinary integer arithmetic on all of them at once. [Daniel's SWAR posts][swar-json] are the best explanation I know. The building block is to repeat a constant across every lane: ```cpp title="Broadcast one byte across a 64-bit word" constexpr uint64_t kOnes = 0x0101010101010101ULL; constexpr uint64_t kHigh = 0x8080808080808080ULL; constexpr uint64_t splat(uint8_t x) { return kOnes * x; } ``` `splat('a')` is `0x6161616161616161`. `kHigh` is a mask of the top bit of every byte. That high bit is our per-lane boolean. `0x80` means this byte failed. `0x00` means it did not. If you have 8 bytes in a 64-bit word `x`, computing `x - splat(32)` subtracts 32 from each byte. It works well if each byte is greater than or equal to 32. Otherwise the operation overflows: if the least significant byte is too small, its most significant bit is set, and you cannot rely on the other lanes. That is why we first keep only ASCII bytes (`x & kHigh` is zero when every byte is below 128), and then apply both bounds. ```cpp title="Per-byte comparisons inside a uint64_t" // High bit set in each byte of x that is strictly less than n. // Requires n <= 128. constexpr uint64_t bytes_less_than(uint64_t x, uint8_t n) { return (x - splat(n)) & ~x & kHigh; } // High bit set in each byte of x that is strictly greater than n. // Requires n < 128. constexpr uint64_t bytes_greater_than(uint64_t x, uint8_t n) { return ((x + splat(127 - n)) | x) & kHigh; } bool eight_bytes_are_lowercase(uint64_t word) { const uint64_t non_ascii = word & kHigh; const uint64_t too_small = bytes_less_than(word, 'a'); const uint64_t too_large = bytes_greater_than(word, 'z'); return (non_ascii | too_small | too_large) == 0; } ``` `bytes_less_than` subtracts `n` from each byte. A lane that was already below `n` underflows, and the `& ~x & kHigh` filter keeps only those underflows. `bytes_greater_than` adds `127 - n`. A lane above `n` crosses 127 and lights the high bit. If `word` is eight lowercase letters, the three masks are zero. If one lane is `'A'` or `'{'` or `0xC3`, the corresponding high bit in the OR is set. The string-level version ORs the failure masks into an accumulator and finishes the leftover bytes with the scalar wraparound test. That's slightly more than one operation per input byte in the main loop, which is the same claim Daniel makes for his JSON-escapable SWAR check. ```cpp title="SWAR lowercase scan with a scalar tail" bool is_ascii_lowercase(std::string_view input) { const auto* p = reinterpret_cast(input.data()); size_t n = input.size(); uint64_t bad = 0; while (n >= 8) { uint64_t word; std::memcpy(&word, p, sizeof(word)); bad |= word & kHigh; bad |= bytes_less_than(word, 'a'); bad |= bytes_greater_than(word, 'z'); p += 8; n -= 8; } while (n--) { bad |= static_cast( static_cast(*p++ - 'a') > 25); } return bad == 0; } ``` The inner loop never branches on a character value. It always consumes eight bytes, always updates `bad`, and always continues. The leftover `n < 8` bytes are not worth the SWAR setup. SWAR is also a good way to understand the SIMD code. NEON is the same algorithm with 16-byte registers and real compare instructions. ## The same check with NEON [ARM NEON][neon] can process 16 bytes at a time. One load, one compare, one OR. There is no SWAR carry trick, because the hardware already knows the lanes are independent. For the most part, your computer is either an ARM machine supporting at least NEON, or an x64 machine supporting at least SSE2. It is easy to distinguish at compile time. A good general strategy, which Daniel uses in the [escaping post][escaping], is to load the data in blocks of 16 bytes and do a few comparisons over those 16 bytes. What about fewer than 16 characters? If you do not want to read past the string, fall back on one of the conventional functions. If the leftover is non-empty but the string was already at least 16 bytes, you can reload the last 16 bytes of the input. That overlapping load avoids a scalar tail. The first NEON version I wrote used two compares, an AND, and a NOT: `vcgeq` against `'a'`, `vcleq` against `'z'`. That works. It is also more work than we need. The wraparound test is faster here too. Unsigned subtract wraps the same way in every lane, so we get `c - 'a'` for 16 bytes at once, then one unsigned compare against 25. ```cpp title="NEON wraparound scan, 16 bytes per iteration" #include bool is_ascii_lowercase(std::string_view input) { if (input.size() < 16) { unsigned errors = 0; for (unsigned char c : input) { errors |= static_cast( static_cast(c - 'a') > 25); } return errors == 0; } const auto* p = reinterpret_cast(input.data()); size_t n = input.size(); uint8x16_t errors = vdupq_n_u8(0); const uint8x16_t va = vdupq_n_u8('a'); const uint8x16_t limit = vdupq_n_u8(25); size_t i = 0; for (; i + 15 < n; i += 16) { const uint8x16_t v = vld1q_u8(p + i); const uint8x16_t d = vsubq_u8(v, va); errors = vorrq_u8(errors, vcgtq_u8(d, limit)); } if (i < n) { const uint8x16_t v = vld1q_u8(p + n - 16); const uint8x16_t d = vsubq_u8(v, va); errors = vorrq_u8(errors, vcgtq_u8(d, limit)); } return vmaxvq_u8(errors) == 0; } ``` That is load, subtract, compare, OR. The two-bound version was load, compare, compare, AND, NOT, OR. Same answer, fewer instructions. Here is what each intrinsic is doing: - `vdupq_n_u8('a')` is the NEON splat. Every lane holds `'a'`. - `vld1q_u8` loads 16 consecutive bytes. The instruction accepts unaligned addresses, so you do not need `memcpy`. - `vsubq_u8` subtracts `'a'` from every lane. Bytes below `'a'` wrap, exactly like `unsigned char(c - 'a')`. - `vcgtq_u8` is an unsigned greater-than. A lane becomes `0xFF` when that wrapped value is greater than 25. - `vorrq_u8` into `errors` is the same OR-accumulation as before. - `vmaxvq_u8` (AArch64) reduces the vector. If any lane is non-zero, some byte was out of range. If the leftover is non-empty but the string was already at least 16 bytes, we reload the last 16 bytes. That overlapping load avoids a scalar tail. Duplicate bytes are fine: we only OR errors. On x64 the same wraparound is subtract plus saturating unsigned subtract. `_mm_subs_epu8(d, 25)` becomes zero when `d <= 25`, and non-zero otherwise. Signed SIMD compares (`_mm_cmplt_epi8`) are a trap for bytes above 127, because those bytes look negative. The wraparound path stays unsigned the whole way. ```cpp title="AVX2 wraparound scan, 32 bytes per iteration" bool is_ascii_lowercase(std::string_view input) { if (input.size() < 32) { unsigned errors = 0; for (unsigned char c : input) { errors |= static_cast( static_cast(c - 'a') > 25); } return errors == 0; } const auto* p = reinterpret_cast(input.data()); size_t n = input.size(); __m256i errors = _mm256_setzero_si256(); const __m256i va = _mm256_set1_epi8('a'); const __m256i limit = _mm256_set1_epi8(25); size_t i = 0; for (; i + 31 < n; i += 32) { const __m256i v = _mm256_loadu_si256(reinterpret_cast(p + i)); const __m256i d = _mm256_sub_epi8(v, va); errors = _mm256_or_si256(errors, _mm256_subs_epu8(d, limit)); } if (i < n) { const __m256i v = _mm256_loadu_si256(reinterpret_cast(p + n - 32)); const __m256i d = _mm256_sub_epi8(v, va); errors = _mm256_or_si256(errors, _mm256_subs_epu8(d, limit)); } return _mm256_testz_si256(errors, errors) != 0; } ``` I ran these on an Intel Xeon with GCC 13.3 (`-O3 -march=native`), scanning a 1 MiB all-lowercase buffer. The happy path is the interesting one: you almost always scan the whole string. | Approach | Throughput | | --- | ---: | | Branchy early `return` | 4.0 GB/s | | 256-byte table | 3.1 GB/s | | Wraparound (GCC autovectorized) | 14.2 GB/s | | SSE2 two-compare | 31.6 GB/s | | SWAR | 53.2 GB/s | | SSE2 wraparound | 56.0 GB/s | | AVX2 wraparound | 63.9 GB/s | The table loses to the naive loop. It is bound by load latency, which is the same thing Daniel saw when he compared a 256-byte identifier table to NEON. GCC already turns the scalar wraparound loop into SIMD, and that is 3.5 times the branchy version, but the hand-written SWAR and AVX2 paths still win by a lot. If the first bad byte is near the start, the branchy loop wins, and the GB/s number becomes meaningless because you barely touch the buffer. That is the only case where I would keep the early `return`. For messier classes, the 256-byte table fails in SIMD: there is no 256-byte gather on NEON. The usual trick is [vectorized classification][neon-ids] (see Langdale and Lemire, [Parsing Gigabytes of JSON per Second][pgjson]). We use the same idea in Ada. Split each byte into two 4-bit nibbles and look those up in two 16-byte tables: ```cpp title="Nibble lookup: a 256-byte class table in two 16-byte vectors" uint8x16_t classify(uint8x16_t input, uint8x16_t table_lo, uint8x16_t table_hi) { const uint8x16_t lo = vandq_u8(input, vdupq_n_u8(0x0F)); const uint8x16_t hi = vshrq_n_u8(input, 4); return vandq_u8(vqtbl1q_u8(table_lo, lo), vqtbl1q_u8(table_hi, hi)); } ``` `vqtbl1q_u8` is a 16-entry table lookup done on all sixteen lanes at once. The low nibble picks a row, the high nibble picks a row, and the AND is 1 only when both halves of the original 256-entry table would have said yes. The table now lives in registers and you classify 16 bytes per call. On x86 the same ideas map to SSE/AVX (`_mm_cmpeq_epi8`, `_mm_shuffle_epi8`). The NEON names change. The loop shape does not. ## Which one should you use? It depends on the check and on the length of the input. For a single interval like `a-z`, the numbers above are the whole story. 1. **Early `return false`** when invalid input is common and usually fails in the first few bytes. That is the only case where the obvious loop is the fast one. 2. **Wraparound** (`unsigned char(c - 'a') <= 25`) for a single interval. Write it as a branchless scalar loop first. GCC and LLVM will often autovectorize it. If that is already off the profile, stop. 3. **SWAR or SIMD wraparound** when the strings are long and you still see the scan in a profile. Prefer subtract-and-compare over two bounds. AVX2 or NEON will beat a portable SWAR loop when you have them. SWAR is the right fallback when you do not. 4. **A 256-byte table** when the check is a union of ranges or a mix of punctuation: hex, unreserved, forbidden host bytes. Size it at 256, not 255. Do not use a table for `a-z`. You are paying a load for a subtract. 5. **Nibble tables in NEON/SSSE3** when that messy class is also the hot SIMD path. That is vectorized classification, the same idea we use in Ada. You can still go further. Unrolling the NEON or AVX2 loop to 32 or 64 bytes hides some of the reduction latency. If invalid input is common but not always at byte 0, check the accumulator every few vectors and return early. On newer ARM, SVE2 `match` / `nmatch` can replace a pile of equality tests for small character sets. I would not start there. Compilers will not invent a 256-byte character class for you, and they will not write the SWAR masks. Measure the version you are about to delete. I used "is this string ASCII lowercase?" as the running example because the check is small enough to see every transformation. In practice, the checks I care about are the other ones: hex, unreserved, forbidden host bytes, whitespace. Those are the loops that show up in a URL parser, and they are often where a simple `for` loop with an early `return` becomes the bottleneck. So far every byte stood on its own. Suppose you want a harder problem. ## A harder problem: is this valid UTF-16? Modern-day text in software can be expected to be Unicode. Unicode is stored in two formats: UTF-8 and UTF-16. UTF-16 is used by several platforms to represent Unicode characters. Microsoft Windows uses it for file names and registry keys. Java and JavaScript use it for strings. UTF-16 represents each character by one or two 16-bit code units. For characters in the Basic Multilingual Plane, which includes most commonly used characters from around the world, a single 16-bit unit suffices. For characters beyond this plane, UTF-16 uses a pair of 16-bit units known as a surrogate pair. That is how it covers a bit more than a million code points while keeping most characters in 16 bits. Values making up surrogate pairs are either high surrogates (`U+D800` to `U+DBFF`) or low surrogates (`U+DC00` to `U+DFFF`). A pair is always made of a high surrogate followed by a low surrogate. Otherwise, we have an error. [Daniel has a post][utf16-neon] on putting a replacement character (`U+FFFD`) wherever that rule breaks. [simdutf][simdutf] is the production version, and it is what V8 uses for `String.toWellFormed`. I only need the boolean: is this valid UTF-16? A basic C++ function might look as follows. ```cpp title="A basic UTF-16 validator" bool is_high_surrogate(char16_t c) { return (c >= 0xD800 && c <= 0xDBFF); } bool is_low_surrogate(char16_t c) { return (c >= 0xDC00 && c <= 0xDFFF); } bool is_valid_utf16(std::u16string_view input) { for (size_t i = 0; i < input.size(); ++i) { if (is_high_surrogate(input[i])) { if (i + 1 < input.size() && is_low_surrogate(input[i + 1])) { ++i; } else { return false; } } else if (is_low_surrogate(input[i])) { return false; } } return true; } ``` The function scans a buffer of `char16_t` values. If a high surrogate is followed by a low surrogate, we skip both. If a high surrogate is not followed by a low surrogate, or if a low surrogate appears without a preceding high surrogate, we return `false`. If the loop finishes, we return `true`. The function should be reasonably efficient. Most of our processors have instructions that process eight 16-bit words per register. Most mobile processors today are 64-bit ARM with NEON. And most UTF-16 text never leaves the BMP. I suspect that this is the typical case: there are relatively few surrogate pairs in most text. If you do not have NEON, you can still ask whether four code units contain any surrogate at all. Mask each 16-bit lane of a `uint64_t` with `0xF800` and look for `0xD800`. No match means those four units are valid BMP text and you can skip the pairing. That is the same SWAR trick as the lowercase scan, just on 16-bit lanes. We can write a function targeting ARM NEON using intrinsic functions. These give us low-level access to NEON. There are comparable intrinsics for Intel/AMD, RISC-V, and so on. The classification inside the loop is the wraparound test from earlier. Adding `0x2800` to a high surrogate wraps it into `0x0000` to `0x03FF`. Adding `0x2400` does the same for a low surrogate. ```cpp title="NEON: validate eight code units at a time" bool is_valid_utf16_neon(std::u16string_view input) { const char16_t* buffer = input.data(); const size_t length = input.size(); const size_t vec_size = 8; size_t i = 0; if (length >= vec_size) { uint16x8_t previous_high_surrogate_mask = vdupq_n_u16(0); for (; i + vec_size <= length; i += vec_size) { const uint16x8_t vec = vld1q_u16( reinterpret_cast(buffer + i)); const uint16x8_t low_surrogate_mask = vcleq_u16( vaddq_u16(vec, vdupq_n_u16(0x2400)), vdupq_n_u16(0x03FF)); const uint16x8_t high_surrogate_mask = vcleq_u16( vaddq_u16(vec, vdupq_n_u16(0x2800)), vdupq_n_u16(0x03FF)); const uint16x8_t offset_high_surrogate_mask = vextq_u16( previous_high_surrogate_mask, high_surrogate_mask, 7); const uint16x8_t offset_low_surrogate_mask = (i + vec_size < length && is_low_surrogate(buffer[i + vec_size])) ? vextq_u16(low_surrogate_mask, vdupq_n_u16(0xFFFF), 1) : vextq_u16(low_surrogate_mask, vdupq_n_u16(0), 1); const uint16x8_t low_not_preceded_by_high = vbicq_u16(low_surrogate_mask, offset_high_surrogate_mask); const uint16x8_t high_not_followed_by_low = vbicq_u16(high_surrogate_mask, offset_low_surrogate_mask); if (vmaxvq_u16(vorrq_u16(low_not_preceded_by_high, high_not_followed_by_low)) != 0) { return false; } previous_high_surrogate_mask = high_surrogate_mask; } } if (i > 0 && is_high_surrogate(buffer[i - 1]) && i < length && is_low_surrogate(buffer[i])) { ++i; } for (; i < length; ++i) { if (is_high_surrogate(buffer[i])) { if (i + 1 < length && is_low_surrogate(buffer[i + 1])) { ++i; } else { return false; } } else if (is_low_surrogate(buffer[i])) { return false; } } return true; } ``` This function uses NEON to validate UTF-16 in chunks of eight code units. For each chunk it loads the data, builds a high surrogate mask and a low surrogate mask, shifts those masks to check for valid pairs across vector boundaries, and returns `false` on a lone high or a lone low. After as many full chunks as we can, the remaining units go through the scalar function. If the last vector unit was a high surrogate and the first tail unit is its low partner, we skip that low so the scalar loop does not treat it as unmatched. Though reasonably efficient, I expect that it is possible to do much better than this function. A reader of Daniel's post proposed a faster alternative that uses the fact that ARM NEON has interleaved loads. When we load the data, we put the most significant bytes in one register and the least significant bytes in the other. The most significant bytes are sufficient to check for errors, so we can check 32 bytes of input by validating just one 16-byte register. ```cpp title="Faster NEON: classify 16 code units from the high bytes" bool is_valid_utf16_neon_v2(std::u16string_view input) { const char16_t* buffer = input.data(); const size_t length = input.size(); const int high_vec = 1; const size_t vec_size = 32; size_t i = 0; if (length * 2 >= vec_size) { uint8x16_t previous_high_surrogate_mask = vdupq_n_u8(0); const uint8_t* buffer8 = reinterpret_cast(buffer); for (; i + vec_size < length * 2; i += vec_size) { const uint8x16x2_t pair = vld2q_u8(buffer8 + i); const uint8x16_t vec = vshrq_n_u8(pair.val[high_vec], 2); const uint8x16_t low_surrogate_mask = vceqq_u8(vec, vdupq_n_u8(0x37)); const uint8x16_t high_surrogate_mask = vceqq_u8(vec, vdupq_n_u8(0x36)); const uint8x16_t offset_high_surrogate_mask = vextq_u8( previous_high_surrogate_mask, high_surrogate_mask, 15); const uint8_t next_char_type = buffer8[i + vec_size + high_vec] >> 2; const uint8x16_t offset_low_surrogate_mask = vextq_u8( low_surrogate_mask, vdupq_n_u8(next_char_type == 0x37 ? 0xFF : 0), 1); const uint8x16_t low_not_preceded_by_high = vbicq_u8(low_surrogate_mask, offset_high_surrogate_mask); const uint8x16_t high_not_followed_by_low = vbicq_u8(high_surrogate_mask, offset_low_surrogate_mask); if (vmaxvq_u8(vorrq_u8(low_not_preceded_by_high, high_not_followed_by_low)) != 0) { return false; } previous_high_surrogate_mask = high_surrogate_mask; } i >>= 1; } if (i > 0 && is_high_surrogate(buffer[i - 1]) && i < length && is_low_surrogate(buffer[i])) { ++i; } for (; i < length; ++i) { const uint16_t surrogate_type = static_cast(buffer[i]) >> 10; if (surrogate_type == 0x36) { if (i + 1 < length && is_low_surrogate(buffer[i + 1])) { ++i; } else { return false; } } else if (surrogate_type == 0x37) { return false; } } return true; } ``` `0x36` is a high surrogate's top six bits (`0xD800 >> 10`). `0x37` is a low surrogate (`0xDC00 >> 10`). [The follow-up paper][utf16-paper] pushes the same idea to 64-unit blocks. You can do even better if you assume that the input rarely contains invalid characters, or rarely contains surrogates at all. I am going to leave that as an exercise for the reader. To benchmark these functions, Daniel used a single string made of 10 million space characters. It is the easiest case: no replacement and no surrogate pairs. I suspect that it also represents a typical case. Using LLVM 16 and an Apple M2, he got: | Approach | Throughput | | --- | ---: | | Regular C | 1.7 GB/s | | NEON | 5.5 GB/s | | Fast NEON | 13 GB/s | So the fast ARM NEON code is about 8 times faster than the conventional code. He was measuring the correction version. The boolean one is the same pairing, with a `return false` instead of a store of `U+FFFD`. [ada]: https://github.com/ada-url/ada [lemire]: https://lemire.me [escaping]: https://lemire.me/blog/2024/05/31/quickly-checking-whether-a-string-needs-escaping/ [tables]: https://lemire.me/blog/2023/11/07/generating-arrays-at-compile-time-in-c-with-lambdas/ [swar]: https://lemire.me/blog/2022/01/21/swar-explained-parsing-eight-digits/ [swar-json]: https://lemire.me/blog/2025/04/13/detect-control-characters-quotes-and-backslashes-efficiently-using-swar/ [neon-ids]: https://lemire.me/blog/2023/09/04/locating-identifiers-quickly-arm-neon-edition/ [pgjson]: https://arxiv.org/abs/1902.08318 [neon]: https://developer.arm.com/architectures/instruction-sets/intrinsics/ [utf16-neon]: https://lemire.me/blog/2024/12/29/efficient-in-place-utf-16-unicode-correction-with-arm-neon/ [simdutf]: https://github.com/simdutf/simdutf [utf16-paper]: https://arxiv.org/abs/2601.06349 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/eliminating-branches-in-cpp-loops Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: "Announcing Ada v4: Validating 35.6M URLs per second" description: "Ada 4.0 is faster on the common path, ships a much smaller binary, and hardens URL parsing with max-length limits, ABI soname 4, and a wave of correctness fixes." date: 2026-07-27 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/release-of-ada-v4" markdown: "https://www.yagiz.co/release-of-ada-v4.md" --- # Announcing Ada v4: Validating 35.6M URLs per second > Ada 4.0 is faster on the common path, ships a much smaller binary, and hardens URL parsing with max-length limits, ABI soname 4, and a wave of correctness fixes. *Published: 2026-07-27 · Tag: performance* --- Ada is the fastest WHATWG-compliant URL parser in the world. It powers URL handling in Node.js, Cloudflare Workers, Redpanda, Kong, Telegram, Datadog, ClickHouse, and more. We are announcing **Ada v4.0.0** — the first major release since [v3.4.4][v344] (23 March 2026). Since that release, [main picked up 79 commits][diff]. This release is about three things: speed on the common path, a materially smaller IDNA/binary footprint, and hardening against incorrect and unsafe inputs. Shared-library consumers should plan for the soname bump: `libada.so.3` → `libada.so.4`. All numbers below were measured on an **Apple M5 Max** (macOS 26.5), Release builds (`-DCMAKE_BUILD_TYPE=Release -DADA_BENCHMARKS=ON`), five Google Benchmark repetitions (means). Same machine, same flags, same datasets for v3.4.4 and current `main` (`16a57723`). ## Performance ### Large realistic dataset (`benchdata`, ~100k URLs) | Benchmark | v3.4.4 | v4.0.0 | Speedup | | --- | ---: | ---: | ---: | | `ada::url` parse + href | 159.3 ns/URL | 128.8 ns/URL | **1.24×** (−19%) | | `url_aggregator` parse + href | 98.5 ns/URL | 81.5 ns/URL | **1.21×** (−17%) | | `ada::can_parse` | 58.5 ns/URL | 28.1 ns/URL | **2.08×** (−52%) | Throughput on this corpus is about **6.3M → 7.8M URLs/s** for `ada::url`, **10.2M → 12.3M URLs/s** for `url_aggregator`, and **17.1M → 35.6M URLs/s** for `can_parse`. ### Small “top sites” set (`bench`) | Benchmark | v3.4.4 | v4.0.0 | Speedup | | --- | ---: | ---: | ---: | | `ada::url` | 164.3 ns/URL | 110.2 ns/URL | **1.49×** (−33%) | | `url_aggregator` | 98.2 ns/URL | 74.6 ns/URL | **1.32×** (−24%) | | `ada::can_parse` | 64.2 ns/URL | 39.9 ns/URL | **1.61×** (−38%) | ### URLSearchParams and hosts | Benchmark | v3.4.4 | v4.0.0 | Speedup | | --- | ---: | ---: | ---: | | `url_search_params` (Indeed fixtures) | 6.40 µs | 3.60 µs | **1.78×** (−44%) | | IPv4 non-decimal (`ada::url`) | 79.8 ns | 74.2 ns | **1.08×** (−7%) | | DNS-style hosts (`ada::url`) | 150.3 ns | 127.6 ns | **1.18×** (−15%) | The already-tuned pure-decimal IPv4 fast path is roughly unchanged after the shared IP parser refactor. The wins show up on the general IPv4 path, host parsing overall, search-params decoding, absolute `http(s)` parsing, and especially `can_parse`. ### What got faster - **`can_parse` fast path** ([#1106][pr-1106], follow-ups [#1111][pr-1111], [#1118][pr-1118], [#1119][pr-1119], [#1189][pr-1189]): validation without constructing a full URL object, with size-limit and IDNA-expansion awareness so it stays equivalent to `parse(...).has_value()`. - **Absolute `http`/`https` parsing** ([#1175][pr-1175]): `try_parse_simple_absolute`, a table-driven single-pass path for already-normalized absolute URLs (clean ASCII host, no credentials/port/IPv4/IDNA, no messy path segments), shared by `url` and `url_aggregator`. Also a faster `get_href` for the common special-scheme case. - **`url_search_params` form-urlencoded decoding** ([#1178][pr-1178]): single-pass `+` / percent decoding, reserved capacity, fewer intermediate strings. - **Shared IPv4/IPv6 parsers** ([#1179][pr-1179]): one implementation for both URL types, hand-rolled number parsers, LUT-backed serialization, less duplication. - **ada-idna** updates through v0.6.0: compressed tables, safer validation, and better throughput on IDNA-heavy hosts. ## Bundle size IDNA table compression is the main story: `src/ada_idna.cpp` dropped from **683 KiB → 400 KiB** (−41%). That flows through to every packaging form we ship. | Artifact (Release, Apple Clang) | v3.4.4 | v4.0.0 | Δ | | --- | ---: | ---: | ---: | | Amalgamated `ada.cpp` | 956 KiB | 706 KiB | **−26%** | | Amalgamated `ada.h` | 404 KiB | 415 KiB | +3% | | `ada.h` + `ada.cpp` zip | 253 KiB | 248 KiB | −2% | | Static `libada.a` | 782 KiB | 618 KiB | **−21%** | | Shared `libada` (stripped `.dylib`) | 616 KiB | 427 KiB | **−31%** | | `ada.cpp.o` `__TEXT` | 579 KiB | 393 KiB | **−32%** | Node.js and other single-header consumers get a noticeably smaller drop-in. Distro shared libraries get a matching soname bump and a smaller `.so`/`.dylib`. ## Breaking changes ### Shared library soname: `libada.so.3` → `libada.so.4` `ADA_LIB_SOVERSION` moved from `3` to `4`. Downstream packages that link the shared library must rebuild against 4.0.0 (or install a package that provides `libada.so.4`). Static linking and the amalgamated `ada.h` / `ada.cpp` pair are unaffected beyond the usual recompile. This follows the ABI discipline we added after [Debian packaging feedback on v3.4.4][issue-1098]: restore exported symbols when needed, keep `abidiff` in CI ([#1099][pr-1099]), and bump the soname when the ABI intentionally changes. ### Configurable maximum URL length Parsing and setters now enforce a configurable maximum on both the raw input and the **normalized href**, including percent-encoding expansion ([#1126][pr-1126]): ```cpp ada::set_max_input_length(2048); // bytes auto url = ada::parse("http://example.com/" + std::string(2048, 'a')); assert(!url); // normalized form too long uint32_t limit = ada::get_max_input_length(); size_t n = url->get_href_size(); // length without allocating ``` The default remains ~4 GB (`UINT32_MAX`), so most applications see no behavior change. If you embed Ada in a service that accepts untrusted URL strings, set a tighter limit. C API: `ada_set_max_input_length` / `ada_get_max_input_length`. ### Correctness fixes that reject previously accepted bad input Several fixes change results for inputs that should never have succeeded. If you were accidentally relying on the old behavior, you will see parse/setter failures instead of silently wrong URLs: - **`set_hostname` / `set_host` no longer keep a partial host** when parsing fails — full rollback ([#1169][pr-1169]), fixing [#1142][issue-1142] (`set_hostname("sneaky.com#legit.org")` and `?` variants). - Hex IPv4 pieces are bounded by **value**, not digit count ([#1185][pr-1185]). - No-scheme state no longer accepts arbitrary input that merely contains a fragment ([#1186][pr-1186]). - URLPattern / IDNA / file-scheme / Windows-drive and related edge cases listed under bug fixes below. ### C API lifetime documentation `ada_string` values from getters are borrowed views into the `ada_url` instance. They are invalidated by any mutating `ada_set_*` / `ada_clear_*` call ([#1091][pr-1091]). This was always true; 4.0.0 documents it explicitly. Prefer `ada_owned_string` getters (e.g. `ada_get_origin`) when you need a stable copy. ## Security and hardening - **URLPattern tokenizer DoS on malformed UTF-8** ([#1125][pr-1125]): malformed sequences always advance the cursor; strict policy errors, lenient policy emits an invalid-char token. - **Max input length** ([#1126][pr-1126]): defense in depth against huge or expansion-amplified URLs. - **Component offset wraparound** ([#1123][pr-1123]): unsigned wraparound eliminated when updating URL component offsets. - **IDNA Bidi rules** ([#1136][pr-1136], [#1145][pr-1145]): LTR labels must end with an allowed class; first-character Bidi rule enforced. - **Empty DNS length check UB** ([#1109][pr-1109]). Fuzzing coverage also grew (serializers harness, stronger invariants, C API harness moved onto the C++ amalgamation). ## Bug reports addressed since v3.4.4 ### Issues | Issue | Summary | | --- | --- | | [#1142][issue-1142] | `set_hostname` silently accepted `#` / `?` and truncated the host — fixed by full rollback on host parse failure ([#1169][pr-1169]). | | [#1127][issue-1127] | Unchecked `simdutf` result in IDNA `to_unicode` — addressed in subsequent ada-idna syncs. | | [#1098][issue-1098] | ABI breakage report from Debian packaging of 3.4.3→3.4.4 — led to restoring `set_scheme` export and adding ABI CI ([#1099][pr-1099]). | ### Spec / correctness fixes (selected) - [#1186][pr-1186] — no-scheme + fragment false accepts - [#1185][pr-1185] — hex IPv4 overflow by digit count - [#1181][pr-1181] — `set_host` dash-dot / port path - [#1167][pr-1167] — hostname canonicalize fast path skipped IPv4/IDNA - [#1166][pr-1166] — explicit port `0` preserved across `set_protocol` on non-special schemes - [#1168][pr-1168] — leading-zero ports in canonicalize - [#1157][pr-1157] — empty host on non-special URL without authority - [#1156][pr-1156] — Windows drive letter must be exactly two code points - [#1139][pr-1139] — `set_protocol` slow-path return consistency - URLPattern: [#1187][pr-1187], [#1188][pr-1188], [#1190][pr-1190], [#1191][pr-1191], [#1171][pr-1171], [#1172][pr-1172], [#1148][pr-1148] Web Platform Tests were rolled forward throughout the cycle. We also listed an official **Kotlin** client alongside the other language bindings ([#1102][pr-1102]). ## Try it ```bash git clone https://github.com/ada-url/ada cd ada cmake -B build -DADA_BENCHMARKS=ON -DCMAKE_BUILD_TYPE=Release cmake --build build ./build/benchmarks/benchdata --benchmark_filter=AdaURL ``` Single-header consumers can amalgamate (`python3 singleheader/amalgamate.py`) or grab the GitHub release assets (`ada.h` / `ada.cpp` / `singleheader.zip`). If you ship Ada as a shared library, plan for **`libada.so.4`**. If you accept untrusted URL strings, set `ada::set_max_input_length` to something appropriate for your service. Thanks to everyone who filed issues, sent fuzz crashes, and opened PRs — especially Debian packaging for the ABI report that sharpened our export/CI story. Full changelog: [v3.4.4...v4.0.0][diff] [v344]: https://github.com/ada-url/ada/releases/tag/v3.4.4 [diff]: https://github.com/ada-url/ada/compare/v3.4.4...main [issue-1098]: https://github.com/ada-url/ada/issues/1098 [issue-1127]: https://github.com/ada-url/ada/issues/1127 [issue-1142]: https://github.com/ada-url/ada/issues/1142 [pr-1091]: https://github.com/ada-url/ada/pull/1091 [pr-1099]: https://github.com/ada-url/ada/pull/1099 [pr-1102]: https://github.com/ada-url/ada/pull/1102 [pr-1106]: https://github.com/ada-url/ada/pull/1106 [pr-1109]: https://github.com/ada-url/ada/pull/1109 [pr-1111]: https://github.com/ada-url/ada/pull/1111 [pr-1118]: https://github.com/ada-url/ada/pull/1118 [pr-1119]: https://github.com/ada-url/ada/pull/1119 [pr-1123]: https://github.com/ada-url/ada/pull/1123 [pr-1125]: https://github.com/ada-url/ada/pull/1125 [pr-1126]: https://github.com/ada-url/ada/pull/1126 [pr-1136]: https://github.com/ada-url/ada/pull/1136 [pr-1139]: https://github.com/ada-url/ada/pull/1139 [pr-1145]: https://github.com/ada-url/ada/pull/1145 [pr-1148]: https://github.com/ada-url/ada/pull/1148 [pr-1156]: https://github.com/ada-url/ada/pull/1156 [pr-1157]: https://github.com/ada-url/ada/pull/1157 [pr-1166]: https://github.com/ada-url/ada/pull/1166 [pr-1167]: https://github.com/ada-url/ada/pull/1167 [pr-1168]: https://github.com/ada-url/ada/pull/1168 [pr-1169]: https://github.com/ada-url/ada/pull/1169 [pr-1171]: https://github.com/ada-url/ada/pull/1171 [pr-1172]: https://github.com/ada-url/ada/pull/1172 [pr-1175]: https://github.com/ada-url/ada/pull/1175 [pr-1178]: https://github.com/ada-url/ada/pull/1178 [pr-1179]: https://github.com/ada-url/ada/pull/1179 [pr-1181]: https://github.com/ada-url/ada/pull/1181 [pr-1185]: https://github.com/ada-url/ada/pull/1185 [pr-1186]: https://github.com/ada-url/ada/pull/1186 [pr-1187]: https://github.com/ada-url/ada/pull/1187 [pr-1188]: https://github.com/ada-url/ada/pull/1188 [pr-1189]: https://github.com/ada-url/ada/pull/1189 [pr-1190]: https://github.com/ada-url/ada/pull/1190 [pr-1191]: https://github.com/ada-url/ada/pull/1191 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/release-of-ada-v4 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: State of URL parsing performance in 2025 description: "Ada is the fastest URL parser, 7.1x faster than cURL for full parsing. At Vercel scale, Ada uses 86% less CPU—saving 34% of one core." date: 2025-12-09 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/state-of-url-parsing-2025" markdown: "https://www.yagiz.co/state-of-url-parsing-2025.md" --- # State of URL parsing performance in 2025 > Ada is the fastest URL parser, 7.1x faster than cURL for full parsing. At Vercel scale, Ada uses 86% less CPU—saving 34% of one core. *Published: 2025-12-09 · Tag: performance* --- ## Introduction Daniel Stenberg (author of cURL) recently [questioned Ada's performance claims](https://bsky.app/profile/bagder.mastodon.social.ap.brid.gy/post/3lymrtbpzezv2), suggesting they were "biased" and "misleading". This post presents updated benchmark results using reproducible methodology on current hardware and software versions. ![Bluesky post](https://www.yagiz.co/content/daniel-stenberg-09-2025.png) [Daniel Lemire](https://lemire.me) and I published our methodology in an academic paper: [Parsing Millions of URLs per Second](https://arxiv.org/abs/2311.10533). The benchmarks below follow the same approach. I've included real-world context using [Vercel's Black Friday data](https://vercel.com/bfcm) to show what these performance differences mean at scale. ## About Ada The current release is Ada v3.3.0, which is used in production by Internet Archive, Node.js, ClickHouse, Redpanda, Kong, Telegram, AdGuard, Datadog, and Cloudflare Workers. Ada v3 introduced URLPattern support alongside its URL parsing capabilities. Ada provides two URL parser implementations optimized for different use cases: `ada::url` optimizes for setters (modifying URL components), while `ada::url_aggregator` optimizes for getters (reading URL components). The benchmarks below test both implementations to show their respective performance characteristics. ## Benchmarks I've run benchmarks on macOS 26 Tahoe 26.1 on my personal Mac Mini 2024 with 14 cores (10 performance and 4 efficiency cores) and 64 GB memory. These benchmarks use Ada commit `662de99d89078a60cc54fa1b696486cfcc814cff`. To reproduce these benchmarks locally, follow these commands: ```bash # Clone the Ada repository git clone https://github.com/ada-url/ada.git cd ada # Checkout the specific commit used in these benchmarks git checkout 662de99d89078a60cc54fa1b696486cfcc814cff # Configure the build with benchmarks enabled cmake -B build \ -DADA_BENCHMARKS=ON \ -DCMAKE_BUILD_TYPE=Release \ -DADA_USE_UNSAFE_STD_REGEX_PROVIDER=ON # Build the project cmake --build build # Run the benchmarks (may require sudo for performance counter access) sudo ./build/benchmarks/benchdata --benchmark_counters_tabular=true ``` This will give you a really nice table with some information about the benchmark and the data before the results such as: ``` loaded db: as4-1 (Apple silicon) Loading /Users/yagiz/coding/ada/ada/ada/build/_deps/url-dataset-src/out.txt Unable to determine clock rate from sysctl: hw.cpufrequency: No such file or directory This does not affect benchmark measurements, only the metadata output. ***WARNING*** Failed to set thread affinity. Estimated CPU frequency may be incorrect. 2025-12-09T10:41:01-05:00 Running ./build/benchmarks/benchdata Run on (14 X 24 MHz CPU s) CPU Caches: L1 Data 64 KiB L1 Instruction 128 KiB L2 Unified 4096 KiB (x14) Load Average: 1.67, 2.03, 3.13 ada spec: Ada follows whatwg/url bad urls: --------------------- ada---count of bad URLs 26 servo/url---count of bad URLs 26 whatwg---count of bad URLs 26 curl---count of bad URLs 130 ------------------------------- bytes/URL: 86.859205 curl spec: Curl follows RFC3986, not whatwg/url curl version : 8.7.1 input bytes: 8688092 number of URLs: 100025 performance counters: Enabled rust version : 1.91.1 zuri : OMITTED ``` ## Defaults on macOS 26 Tahoe macOS 26 Tahoe comes with cURL 8.7.1. Using the default version of cURL, the performance of all URL parsers are as follows: ``` ➜ sudo ./build/benchmarks/benchdata --benchmark_counters_tabular=true ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- Benchmark Time CPU Iterations GHz cycle/byte cycles/url instructions/byte instructions/cycle instructions/ns instructions/url ns/url speed time/byte time/url url/s ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- BenchData_BasicBench_AdaURL_href 15691123 ns 15670386 ns 44 4.28376 7.50021 651.462 36.7012 4.89336 20.962 3.18784k 152.077 554.427M/s 1.80366ns 156.665ns 6.38306M/s BenchData_BasicBench_AdaURL_aggregator_href 10755944 ns 10742391 ns 64 4.35728 5.13727 446.219 26.5566 5.16939 22.5245 2.30668k 102.408 808.767M/s 1.23645ns 107.397ns 9.31124M/s BenchData_BasicBench_AdaURL_CanParse 6515744 ns 6507162 ns 105 4.33009 3.10174 269.414 16.3532 5.27229 22.8295 1.42043k 62.219 1.33516G/s 748.975ps 65.0554ns 15.3715M/s BasicBench_whatwg 30237078 ns 30214870 ns 23 4.17769 14.328 1.24452k 79.8836 5.57535 23.2921 6.93862k 297.896 287.544M/s 3.47773ns 302.073ns 3.31046M/s BasicBench_CURL 76373611 ns 76343111 ns 9 4.19096 35.5292 3.08604k 190.592 5.36437 22.4819 16.5547k 736.356 113.803M/s 8.7871ns 763.24ns 1.3102M/s BasicBench_ServoUrl 26764394 ns 26734192 ns 26 4.22544 12.858 1.11683k 65.8217 5.11914 21.6306 5.71722k 264.312 324.981M/s 3.07711ns 267.275ns 3.74146M/s ``` Ada is the fastest URL parser in this comparison. The `CanParse` method (which only validates URL correctness without parsing all components) completes each validation in 65 nanoseconds and processes 15.37 million URLs per second, which is **11.7x faster than CURL**. The `aggregator_href` variant (full parsing with all URL components) takes 107 nanoseconds per URL and handles 9.31 million URLs per second, which is **7.1x faster than CURL**. Ada maintains full WHATWG URL Standard compliance. CURL follows RFC 3986, which is a different specification. **Important Note on Fairness**: This comparison isn't entirely apples-to-apples. RFC 3986 is significantly simpler and can be implemented without extra memory allocations, whereas the WHATWG URL Standard is considerably more complex due to the numerous edge cases it handles (internationalized domain names, special schemes, percent encoding normalization, etc.). **Ada is doing substantially more work per URL** while still achieving better performance. The performance comparison would be more direct if both parsers implemented the same specification. ## What This Means in Practice To understand what these numbers mean for production systems, I looked at Vercel's Black Friday/Cyber Monday 2025 traffic. During this period, Vercel handled **115.8 billion requests** over four days, with peak traffic reaching **518,027 requests per second**. If each request requires parsing a URL (which is common in web servers for routing, logging, and request handling), here's what the CPU requirements would look like using full URL parsing: **Using cURL:** - Each core handles 1.31 million URLs per second - Peak traffic of 518,027 req/s requires **39.5% of one CPU core** just for URL parsing - That's 518,027 ÷ 1,310,000 = 0.395 cores **Using Ada url_aggregator:** - Each core handles 9.31 million URLs per second - Same peak traffic requires only **5.6% of one CPU core** - That's 518,027 ÷ 9,310,000 = 0.056 cores **The difference: 33.9% of one core saved, an 86% reduction in compute resources dedicated to URL parsing.** ## Cost Impact at Scale For a platform processing over 100 billion requests during a peak shopping weekend, this difference translates to meaningful savings: - **Reduced cloud compute costs**: Saving 34% of a core per server compounds across entire fleets. For a deployment with 1,000 servers handling URL-heavy workloads, this translates to 340 cores worth of compute capacity freed up for other tasks. - **Lower power consumption**: Less CPU utilization means reduced electricity costs in data centers. Even a fraction of a core, when multiplied across thousands of machines running 24/7, creates measurable energy savings. - **Smaller cooling requirements**: Reduced CPU load generates less heat, lowering cooling infrastructure demands. - **Higher capacity on existing hardware**: The freed-up compute headroom provides a buffer for traffic spikes without requiring additional provisioning. While 34% of one core might seem small in isolation, at cloud scale these efficiency gains compound significantly. For services processing billions of URLs daily across distributed infrastructure, even fractional improvements per machine translate to substantial cost reductions and improved resource utilization across the entire fleet. ## Updating to latest version of cURL The benchmarks above used cURL 8.7.1, which ships by default with macOS 26 Tahoe. I also tested against the latest version of cURL (8.17.0) available via Homebrew. The latest version performs at 892 nanoseconds per URL and processes 1.12 million URLs per second, which is slightly slower than the default version (763ns, 1.31M URLs/s). Using the latest cURL version with Ada's `url_aggregator`, the CPU core difference would be even larger—**46.2% of one core with latest cURL versus 5.7% with Ada, an 88% reduction**. The performance regression between cURL 8.7.1 and 8.17.0 is notable. The latest version takes **129 nanoseconds longer per URL** (892ns vs 763ns), which represents a **17% slowdown**. At Vercel's peak traffic scale of 518,027 requests per second, this difference would require **6.7% more of a CPU core** just to maintain the same throughput. This shows how even minor performance regressions in parsing libraries can have measurable infrastructure cost implications at scale. To run the benchmarks with a specific cURL version (tested with v8.17.0), use these commands: ```bash # Install the latest cURL via Homebrew (macOS) brew install curl # Clone and configure with specific cURL version git clone https://github.com/ada-url/ada.git cd ada cmake -B build \ -DADA_BENCHMARKS=ON \ -DCMAKE_BUILD_TYPE=Release \ -DADA_USE_UNSAFE_STD_REGEX_PROVIDER=ON \ -DCURL_INCLUDE_DIR=/opt/homebrew/opt/curl/include \ -DCURL_LIBRARY=/opt/homebrew/opt/curl/lib/libcurl.dylib \ -DCMAKE_FIND_FRAMEWORK=LAST \ -DCMAKE_FIND_APPBUNDLE=LAST # Build and run cmake --build build sudo ./build/benchmarks/benchdata --benchmark_counters_tabular=true ``` And here are the results: ``` ➜ sudo ./build/benchmarks/benchdata --benchmark_counters_tabular=true ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- Benchmark Time CPU Iterations GHz cycle/byte cycles/url instructions/byte instructions/cycle instructions/ns instructions/url ns/url speed time/byte time/url url/s ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- BenchData_BasicBench_AdaURL_href 15803060 ns 15782000 ns 44 4.2983 7.57612 658.055 36.6275 4.8346 20.7806 3.18144k 153.097 550.506M/s 1.81651ns 157.781ns 6.33792M/s BenchData_BasicBench_AdaURL_aggregator_href 10969095 ns 10968484 ns 64 4.24262 5.24663 455.719 26.5629 5.06284 21.4797 2.30723k 107.414 792.096M/s 1.26247ns 109.657ns 9.11931M/s BenchData_BasicBench_AdaURL_CanParse 6682434 ns 6676443 ns 106 4.11577 3.13855 272.612 16.3542 5.21076 21.4463 1.42052k 66.2359 1.30131G/s 768.459ps 66.7477ns 14.9818M/s BasicBench_whatwg 31119821 ns 31092348 ns 23 4.15052 14.6009 1.26822k 79.8787 5.47082 22.7067 6.9382k 305.557 279.429M/s 3.57873ns 310.846ns 3.21703M/s BasicBench_CURL 89264557 ns 89178750 ns 8 4.1492 42.1106 3.6577k 233.565 5.54647 23.0134 20.2873k 881.542 97.4233M/s 10.2645ns 891.565ns 1.12162M/s BasicBench_ServoUrl 27290262 ns 27273400 ns 25 4.17501 12.8397 1.11524k 65.8058 5.1252 21.3977 5.71584k 267.124 318.556M/s 3.13917ns 272.666ns 3.66749M/s ``` ## Conclusion This benchmark demonstrates that Ada delivers significant performance advantages over cURL for URL parsing, with **7.1x faster full parsing** and **86% reduction in CPU utilization** at scale. More importantly, it shows that URL parsing remains highly efficient even at extreme scale—requiring less than half a CPU core for over half a million requests per second. The performance difference matters most when multiplied across large infrastructure deployments. While a single server saves only a fraction of a core, this compounds to meaningful cost savings and improved resource utilization across fleets of hundreds or thousands of machines. ### Try Ada Ada is open source and available for integration into your projects: - **GitHub**: [https://github.com/ada-url/ada](https://github.com/ada-url/ada) - **Documentation**: [https://ada-url.com](https://ada-url.com) - **Paper**: [Parsing Millions of URLs per Second](https://arxiv.org/abs/2311.10533) Ada supports C++, C, Node.js, Python, Rust, and other languages through various bindings. It's already used in production by major platforms including Node.js, ClickHouse, Cloudflare Workers, and Datadog. All benchmarks in this post are reproducible using the commands provided. The methodology and results are published in our peer-reviewed academic paper for full transparency. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/state-of-url-parsing-2025 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Release of Ada v3.0 with URLPattern description: "Ada 3.0 adds URLPattern matching for routes, used in Node.js, ClickHouse, and Cloudflare Workers." date: 2025-01-30 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/release-of-ada-v3" markdown: "https://www.yagiz.co/release-of-ada-v3.md" --- # Release of Ada v3.0 with URLPattern > Ada 3.0 adds URLPattern matching for routes, used in Node.js, ClickHouse, and Cloudflare Workers. *Published: 2025-01-30 · Tag: performance* --- Ada is the fastest WHATWG-compliant URL parser in the world. Ada is now used by major products and softwares like Node.js, ClickHouse, Redpanda, Kong, Telegram and Cloudflare Workers. [We are excited][contributors] to announce the release of Ada v3.0. This major release includes a lot of new features, improvements and bug fixes. Since our last release (v2.9.2), published on September 2, 2024, [there have been 262 commits written by 9 contributors][diff-between-292]. ## New features - Ada now supports Bazel build system. - Ada now supports Unicode 15.1.0 through [ada_idna][ada-idna] library. - Ada now includes an experimental URLPattern implementation that allows you to define URL patterns for your routes, and match them. - It doesn't provide a regex engine and leaves the decision of choosing the right engine to the implementor. This is done as a security measure since the default std::regex engine is not safe and open to DDOS attacks. If your url pattern inputs come from an untrusted source (like a user), you *should not* use std::regex in production. Unsafe regular expression libraries have unbounded memory consumption and execution times. We recommend using the "v8" or "Google RE2" regex engines which are safe and performant. - In future releases, we will expand our C API to support URLPattern as well. ```cpp // Define a regex engine that conforms to the following interface // For example, we will use v8 regex engine class v8_regex_provider { public: v8_regex_provider() = default; using regex_type = v8::Global; static std::optional create_instance(std::string_view pattern, bool ignore_case); static std::optional>> regex_search( std::string_view input, const regex_type& pattern); static bool regex_match(std::string_view input, const regex_type& pattern); }; // Define a URLPattern auto pattern = ada::parse_url_pattern("/books/:id(\\d+)", "https://example.com"); // Check validity if (!pattern) { return EXIT_FAILURE; } // Match a URL auto match = pattern->match("https://example.com/books/123"); // Test a URL auto matched = pattern->test("https://example.com/books/123"); ``` ## Changes - Ada library now includes the source code of CPM (C++ Package Manager) to avoid the need of installing it through the internet. ## Breaking changes - Ada now uses C++20 features which requires minimum of GCC 12, LLVM 14 or Visual Studio 2019. - `ada::errors::generic` has been replaced with `ada::errors::type_error` We have made a lot of improvements to the library, and we strongly suggest you to upgrade to v3.0 to benefit for these improvements. Full changelog can be found from [Github][github-release] [ada-idna]: https://github.com/ada-url/idna [contributors]: https://github.com/ada-url/ada/graphs/contributors [diff-between-v292]: https://github.com/ada-url/ada/compare/v2.9.2...v3.0.0 [github-release]: https://github.com/ada-url/ada/releases/tag/v3.0.0 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/release-of-ada-v3 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: "Developing fast & built-in task runner in Node.js core" description: "How Node.js got a built-in task runner, and the steps that made `node --run` fast." date: 2024-06-17 tag: coding author: Yagiz Nizipli canonical: "https://www.yagiz.co/developing-fast-builtin-task-runner" markdown: "https://www.yagiz.co/developing-fast-builtin-task-runner.md" --- # Developing fast & built-in task runner in Node.js core > How Node.js got a built-in task runner, and the steps that made `node --run` fast. *Published: 2024-06-17 · Tag: coding* --- With this blog post, I'm going to explain and analyze the steps I've taken to land a super-fast, built-in task runner in Node.js core. For those who are not familiar with the term "task runner", it's a command-line tool to run a task specified in a `package.json` file in node.js projects. ```json title="Example package.json" { "scripts": { "start": "node index.js" } } ``` ## From a user perspective Whenever, a person executes `npm run start` on this project, it makes a series of executions that makes the operation slow. Here are the steps taken to spawn a process and executes your command: - User types `npm run start` in the terminal. - `npm` tries to find the closest package.json file - It starts from current directory and traverses up to the root of your operating-system, while making `stat` calls in every folder to search for `package.json` file. This means if your current path is `/home/username/projects/my-project`, it will look for `package.json` in the following directories: - `/home/username/projects/my-project` - `/home/username/projects` - `/home/username` - `/home` - `/` - Every parent directory of your current directory gets added to the `PATH` environment variable with a suffix of `node_modules/.bin`. - This is done to make sure that you can have a command like `biome check .` in your `package.json` even though Biome binary is available at `node_modules/.bin` folder. - It reads the `package.json` and looks into `scripts[key]` field to find the command to execute. - It spawns a process and executes in the shell. - This is the slowest part of the operation because it involves spawning a new process and executing the command in the shell. Some package managers like `npm` adds several `npm` specific environment variables into the newly spawned process, such as `npm_lifecycle_event`, `npm_config_user_agent` and `npm_package_json`. ## The problem with `npm` task runner Before diving into the technical implementation, let's analyze the problems with the current task runner in `npm`: - `npm` does not dynamically load it's subcommands like `run`, `install`, `test`, etc. This means that every time you run `npm`, it has to load all the subcommands, even though you are only interested in running a script. - For those who are unfamiliar with `node` internals, whenever you `require()` a module in a project, it makes a filesystem call to the operating system to read the file and parse it. This gets cached on common.js applications (but not on ESM), but regardless separating and having multipel files results in slower startup times. - This will likely be optimized in the future, but for the time being it's a problem that effects the execution time. - `npm` runs in the context of Node.js, which means it has to load the Node.js runtime and execute the command in the shell. - By default, in order to execute a shell command, you have to initialize a Node.js process, that initializes V8 engine, parses the `npm` JS code, and then executes the command in the shell. This is a slow process. (Even writing this sentence is slow, imagine how slow it is to execute a shell command on a Node.js library that is not optimized for this purpose.) ## Solution The solution to the problem is to create a new command in Node.js that is optimized for running tasks in a project. This command should be able to run a task specified in a `package.json` file, without having to load the entire Node.js runtime. The original implementation was written in JavaScript to make sure that the Node.js project, contributors and technical steering committee members are open to the idea of having a built-in task runner in Node.js core. After the initial pull-request landed, I've re-written the implementation in C++ to make sure that the performance is optimal. There is still some things that needed to be done to reduce the overhead to less than 10ms, but the current implementation is already faster than all alternatives. If you're interested in contributing and optimizing this implementation even more, please let me know. I'm more than happy to help you get started. ## Benchmarks ### JavaScript solution inside Node.js project This implementation is available at [this pull-request][cli-implement-node-run-script-in-package-json]. ``` ❯ hyperfine './out/Release/node run test' 'npm run test' -i Benchmark 1: ./out/Release/node run test Time (mean ± σ): 29.3 ms ± 1.1 ms [User: 23.2 ms, System: 3.1 ms] Range (min … max): 27.6 ms … 33.2 ms 97 runs Warning: Ignoring non-zero exit code. Benchmark 2: npm run test Time (mean ± σ): 185.7 ms ± 9.2 ms [User: 136.7 ms, System: 30.3 ms] Range (min … max): 174.7 ms … 212.9 ms 15 runs Warning: Ignoring non-zero exit code. Summary ./out/Release/node run test ran 6.34 ± 0.40 times faster than npm run test ``` ### C++ re-write of the task runner Almost 10ms faster than the JavaScript solution, which doesn't require V8 to load and execute the original JavaScript implementation. ``` ❯ hyperfine '../node/main-branch --run test' '../node/cpp-rewrite --run test' 'npm run test' -i Benchmark 1: ../node/main-branch --run test Time (mean ± σ): 28.9 ms ± 0.9 ms [User: 24.2 ms, System: 3.4 ms] Range (min … max): 27.5 ms … 31.7 ms 96 runs Warning: Ignoring non-zero exit code. Benchmark 2: ../node/cpp-rewrite --run test Time (mean ± σ): 18.3 ms ± 0.6 ms [User: 16.0 ms, System: 1.5 ms] Range (min … max): 17.5 ms … 20.8 ms 139 runs Warning: Ignoring non-zero exit code. ``` ## Timeline For those interested in the technical implementation, here's a list of series of pull-requests that lead to make `node --run` task runner from scratch to a stable release candidate. In overall, the whole implementation process took almost 3 months (March 22 to June 12). 1. [cli: implement `node --run `][cli-implement-node-run-script-in-package-json] 2. [src: rewrite task runner in c++][src-rewrite-task-runner-in-cpp] 3. [src: fix positional args in task runner][src-fix-positional-args-in-task-runner] 4. [test: add env variable test for --run][test-add-env-variable-test-for-run] 5. [cli: add `NODE_RUN_SCRIPT_NAME` env to `node --run`][cli-add-node-run-script-name-env-variable-to-node-run] 6. [cli: add `NODE_RUN_PACKAGE_JSON_PATH` env][cli-add-node-run-package-json-path-env] 7. [src: traverse parent folders while running `--run`][src-traverse-parent-folders-while-running-run] 8. [doc: move `node --run` stability to release candidate][doc-move-node-run-stability-to-release-candidate] [cli-implement-node-run-script-in-package-json]: https://github.com/nodejs/node/pull/52190 [src-rewrite-task-runner-in-cpp]: https://github.com/nodejs/node/pull/52609 [src-fix-positional-args-in-task-runner]: https://github.com/nodejs/node/pull/52810 [test-add-env-variable-test-for-run]: https://github.com/nodejs/node/pull/52811 [cli-add-node-run-script-name-env-variable-to-node-run]: https://github.com/nodejs/node/pull/53032 [cli-add-node-run-package-json-path-env]: https://github.com/nodejs/node/pull/53058 [src-traverse-parent-folders-while-running-run]: https://github.com/nodejs/node/pull/53154 [doc-move-node-run-stability-to-release-candidate]: https://github.com/nodejs/node/pull/53433 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/developing-fast-builtin-task-runner Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Dear 20 year old Software Engineer description: "A 2015 letter to a 20-year-old engineer: career choices, burnout, and what actually compounds." date: 2024-05-02 tag: personal author: Yagiz Nizipli canonical: "https://www.yagiz.co/dear-20-year-old-software-engineer" markdown: "https://www.yagiz.co/dear-20-year-old-software-engineer.md" --- # Dear 20 year old Software Engineer > A 2015 letter to a 20-year-old engineer: career choices, burnout, and what actually compounds. *Published: 2024-05-02 · Tag: personal* --- > I initially wrote this article for a Medium publication at Mar 21, 2015. Almost 9 years ago! The original blog post is still [ available at Medium](https://medium.com/@anonrig/dear-20-year-old-sofware-engineer-29d7af9a5c1d) Before writing a letter to 20 year old me, it’s better to introduce myself. I am a Computer Engineering student at Sabanci University, Istanbul, Turkey. I’ve worked for more than 5 companies which ranges from startups to corporations. I’ve experienced different environments, I’ve learned different programming languages, but most important of them all, I’ve tried every position available that I can be in within the range of my age. Currently, I am working as a Software Engineer at Signalive and currently developing Snapmail in my part time. > Dear 20 year old me, > > I know you are trying to achieve your best. Within the time you live, you keep improving yourself. Learning new programming languages, testing different environments, developing different mobile applications. > > For what? You know that don’t you? > > You are trying to improve yourself to be better, and finally, to earn more money. It’s true what they say. More money doesn’t bring you happiness. But as the elders say, you have to make a mistake to learn what does it mean to “make a mistake”. Don’t be afraid to make mistakes. > > It’s important to teach. Don’t worry, I don’t mean College. While teaching something, you are going to figure out what things that make you excited and more passionate. > > Do what makes you passionate. I’ve always read different stories in Medium. Computer Engineers was always frustrated by the work load of their jobs. It’s true. Computer Engineering is not a joke. A lot of people are going to depend on you. Not to fail. Not to disappoint. But not only in Computer Industry, but also in life too. > > A lot of friends including loved ones are going to come up with different applications, and going to involve you into it. Don’t hesitate to say no. In long term, saying no to something that doesn’t make you feel excited, will make you happy. > > Try everything. Don’t worry. Your professors are going to say, “this area of computer engineering is not for you” but it is ok. Try that. Don’t hesitate. If you are going to fail, you should fail because of the choices that you make, not because of others. > > Language doesn’t matter. Don’t ever forget that. While learning different programming languages, you are going to love one more than another. Don’t limit yourself with just that. iOS or Android, Native mobile application or hybrid mobile application, they all serve to one purpose. To achieve to your task. > > Try to balance. It is important to balance your love life, friends and family. If you are going to work hard, you should know that one of these 3 portions of your life is going to feel unloved. > > Set small but efficient goals. If you don’t make your goal achievable, it will frustrate you. But most importantly, it will demotivate you. Set small goals in the path to success. > > You have read a lot of stories from Facebook. Stories that involve commonly made mistakes. They all say “Don’t work hard”. Being a workaholic isn’t a good thing. You are going to regret it when you get older. (Trust me, you are going to get older.) > > You have used a lot of open-source projects. A lot. Don’t forget to contribute to them. Contributing to open-source world will improve you more than anything. You will feel more ambitious then ever. Don’t be afraid. It will motivate you, I promise. > > Sincerely, > Yagiz Nizipli --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/dear-20-year-old-software-engineer Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Recap 2023 - The year of hard work and new beginnings description: "A 2023 recap: Ada, Node.js TSC work, talks, and becoming a parent." date: 2023-12-25 tag: personal author: Yagiz Nizipli canonical: "https://www.yagiz.co/recap-2023" markdown: "https://www.yagiz.co/recap-2023.md" --- # Recap 2023 - The year of hard work and new beginnings > A 2023 recap: Ada, Node.js TSC work, talks, and becoming a parent. *Published: 2023-12-25 · Tag: personal* --- It's been a really long year. I've had a lot of ups and downs, but I'm glad I've made it through. I'm looking forward to the next year, and I hope it's a good one. For the people interested about what I've been up to this year, here's a recap of what I've been up to. ## Personal Life 2023 was a really exciting year for me, and my family: - I became a father, and had a daughter named Ada. - We moved to New Jersey and finally abandoned Manhattan and city lifestyle. - We got our green cards, and we're now permanent residents of the United States. ### From Rust to Parenthood: Yagiz Nizipli's Journey Here's an interview I did with [Nearform][nearform] about my journey through Node.js, URL parsing, and personal growth. [![From Rust to Parenthood: Yagiz Nizipli's Journey through Node.js, URL Parsing, and Personal Growth](https://i3.ytimg.com/vi/1ex8dZ7i8_M/maxresdefault.jpg)](https://www.youtube.com/watch?v=1ex8dZ7i8_M) ## Work I've made an exciting career change and started working on developer tooling. This career change made me realize how much I love working on developer tooling, and how much I missed it. I'm really excited about the future, and I'm looking forward to working on developer tooling for the foreseeable future. I've changed jobs this year. Until November 20, I was working at [Noonlight][noonlight] as a Senior Software Engineer. Currently I'm working at [Sentry][sentry] working on error tracking & performance. ## Open-Source This is the year I've spent most of my time coding and contributing to Node.js. Every project I've worked on has been open source, and had the goal of improving the Node.js performance. ![2023 Github Contributions](https://www.yagiz.co/content/2023-github-contributions.png) - Created [4108 commits][github-anonrig] on Github. - I've created [91 pull requests][nodejs-prs] for Node.js. - I've introduced 3 new dependencies to Node.js core: - [Ada URL parser][ada-url-parser], for fast parsing of URLs. - [simdutf][simdutf], for fast UTF-8 validation and transcoding. - [simdjson][simdjson], for fast JSON parsing. - Wrote the fastest URL parser in the world [Ada URL parser][ada-url-parser] with [Daniel Lemire][daniel-lemire] - Ada became the URL parser of Node.js and later got adopted by Cloudflare workers. - Published a paper called [Parsing millions of URLs per second][parsing-millions-of-urls-per-second]. - Created an IDNA implementation for Ada, which got adopted by [Clickhouse][clickhouse-ada-idna]. [Ada IDNA][ada-idna] currently powers `domainToASCII` and `domainToUnicode` in Node.js, and makes `URL` class ICU-free. - I spoke at [NodeConf.eu][nodeconf-eu] with [Daniel Lemire][daniel-lemire] about the Ada URL parser. - I became a voting member of the [Node.js Technical Steering Committee][nodejs-tsc] and [OpenJS Foundation][openjsf-cpc]. - I became the [Node.js Performance Strategic Initiative][performance-strategic-initiative] champion of Node.js and later [resigned][performance-strategic-initiative-leave] from my position. - I attended and spoke at DevTools.fm podcast, OpenJS Collaborator summit, and gave a talk for NearForm's Fireside chats. - I started rewriting [pnpm][pnpm] in Rust called [pacquet][pacquet] and donated it to the pnpm organization. - I joined the [Web-Platform Tests][web-platform-tests] and [pnpm][pnpm-organization] organization. - My JavaScript library [fast-querystring][fast-querystring] got adopted by several large organizations and governments such as [Canada][fast-querystring-canada], [Oracle][fast-querystring-oracle], [Cisco][fast-querystring-cisco] and [Inspire Flex][fast-querystring-inspire-flex]. ## What's next? I've spent 2023 working on Node.js, and giving back to the community. I'm planning to continue working on Node.js, and developer tooling in 2024. Here are my new year's resolutions: ### Work I'll be focusing on improving my Rust skills and improve my expertise in developer tooling and error tracking. Specifically, I'll be working on Sentry's performance and error tracking products. My goal is to improve my knowledge on Rust and C++ performance. I'll try to learn more and apply my findings to Sentry's performance and error tracking products. ### Open-Source 2023 was a year of ups and downs. Especially in the open-source world. I've spent this year creating an initiative and influence over the ecosystem to improve Node.js performance. I've spent a lot of time and energy on this, and I'm glad I've made it through. I'm really proud of what I've accomplished, but I'm also exhausted. I think it's time for me to take a step back from Node.js performance and focus on other things. I'll continue my work on Node.js Technical Steering Committee, but I'll take less responsibility starting by [finding a new chair for the Node.js performance team][finding-new-chair-for-nodejs-performance] bi-weekly meetings. On top of that, I'll be working on a high-performance project similar to [Ada URL][ada-url-parser] with [Daniel Lemire][daniel-lemire]. ## Thank you I'd like to thank Matteo Collina, James Snell, Daniel Lemire, Anna Henningsen, and the entire Node.js community for their support and guidance. I'd also like to thank my wife for her support and patience. And most importantly, I'd like to thank my daughter Ada for being the best thing that happened to me in 2023. [ada-idna]: https://github.com/ada-url/idna [ada-url-parser]: https://github.com/ada-url/ada [clickhouse-ada-idna]: https://github.com/ClickHouse/ClickHouse/pull/57969 [daniel-lemire]: https://lemire.me [fast-querystring]: https://www.npmjs.com/package/fast-querystring [fast-querystring-canada]: https://code.open.canada.ca/en/dependencies.html [fast-querystring-cisco]: https://www.cisco.com/c/dam/en_us/about/doing_business/open_source/docs/CiscoAppDynamicsOrionSolutionManagement-510-1686166002.pdf [fast-querystring-inspire-flex]: https://university.quadient.com/c/portal/documents/find?name=open-source-licenses-inspire [fast-querystring-oracle]: https://docs.oracle.com/en/industries/hospitality/opera-cloud/23.2/ocslg/ch_licensing_information.htm#LicensingInformation-82045BEE [finding-new-chair-for-nodejs-performance]: https://github.com/nodejs/performance/issues/143 [github-anonrig]: https://github.com/anonrig [parsing-millions-of-urls-per-second]: https://onlinelibrary.wiley.com/doi/full/10.1002/spe.3296 [performance-strategic-initiative]: https://github.com/nodejs/node/pull/47424 [performance-strategic-initiative-leave]: https://github.com/nodejs/node/pull/49641 [nearform]: https://nearform.com [nodeconf-eu]: https://www.youtube.com/watch?v=tQ-6OWRDsZg [nodejs-prs]: https://github.com/nodejs/node/pulls?q=is%3Apr+author%3Aanonrig+merged%3A2023-01-01..2024-01-01+is%3Aclosed [nodejs-tsc]: https://github.com/nodejs/tsc [noonlight]: https://noonlight.com [openjsf-cpc]: https://github.com/openjs-foundation/cross-project-council/pull/1207 [pacquet]: https://github.com/pnpm/pacquet [pnpm]: https://pnpm.io [pnpm-organization]: https://github.com/pnpm [sentry]: https://sentry.io [simdutf]: https://github.com/simdutf/simdutf [simdjson]: https://github.com/simdjson/simdjson [web-platform-tests]: https://github.com/web-platform-tests/wpt/issues/43580 --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/recap-2023 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Improving Node.js loader performance description: "Where CommonJS and ESM loaders spend time in Node.js, and what we changed to make them faster." date: 2023-12-12 tag: performance author: Yagiz Nizipli canonical: "https://www.yagiz.co/improving-nodejs-loader-performance" markdown: "https://www.yagiz.co/improving-nodejs-loader-performance.md" --- # Improving Node.js loader performance > Where CommonJS and ESM loaders spend time in Node.js, and what we changed to make them faster. *Published: 2023-12-12 · Tag: performance* --- > This article is originally available at [Sentry.engineering][sentry-engineering-blog]. Node.js supports 2 different modules. EcmaScript and CommonJS modules. ES modules are the official standard for modules in JavaScript and they are supported by all modern browsers. CommonJS modules are the modules that Node.js uses by default. They are not supported by browsers and they are not the official standard. However, they are still widely used. ## How does Node.js load the entry point? In order to differentiate which loader to use, Node.js depends on several factors. The most important one is the file extension. If the file extension is `.mjs`, Node.js will use the ES module loader. If the file extension is `.cjs`, Node.js will use the CommonJS module loader. If the file extension is `.js`, Node.js will use the CommonJS module loader if the `package.json` file has `"type": "commonjs"` field (or simply doesn't have the `type` field). If the `package.json` file has `"type": "module"` field, Node.js will use the ES module loader. This decision is made in `lib/internal/modules/run_main.js` file. You can see a simplified version of [the code][prior-to-optimization-run-main] below: ```js const { readPackageScope } = require('internal/modules/package_json_reader') function shouldUseESMLoader(mainPath) { // Determine the module format of the entry point. if (mainPath && mainPath.endsWith('.mjs')) { return true } if (!mainPath || mainPath.endsWith('.cjs')) { return false } const pkg = readPackageScope(mainPath) switch (pkg.data?.type) { case 'module': return true case 'commonjs': return false default: { // No package.json or no `type` field. return false } } } ``` `readPackageScope` traverses the directory tree upwards until it finds a `package.json` file. Prior to the optimizations done on this post, `readPackageScope` calls an internal version of `fs.readFileSync` until it finds a `package.json` file. This synchronous call makes a filesystem operation and communicates with Node.js C++ layer. This operation has performance bottlenecks depending on the value/type it returns because of the cost of serialization/deserialization of data. This is why we want to avoid calling `readPackage` a.k.a. `fs.readFileSync` inside `readPackageScope` as much as possible. ## How Node.js parses `package.json`? By default, `readPackage` calls an internal version `fs.readFileSync` to read the `package.json` file. This synchronous call returns a string from Node.js C++ layer, which later gets parsed using V8's `JSON.parse()` method. Depending on the validity of this JSON, Node.js checks and creates an object that's required for the remaining of the loaders to perform. These fields are `pkg.name`, `pkg.main`, `pkg.exports`, `pkg.imports` and `pkg.type`. If the JSON has faulty syntax, Node.js will throw an error and exit the process. The output of this function is later cached at an internal `Map` to avoid calling `readPackageScope` again for the same path. This cache is stored for the rest of the process lifetime. ## Usage of `package.json` fields and the reader Before we dive into what optimizations we can do, let's see how Node.js uses these fields. The common use cases in Node.js codebase for parsing and re-using `package.json` fields are: - `pkg.exports` and `pkg.imports` are used to resolve different modules according to your input. - `pkg.main` is used to resolve the entry point of the application. - `pkg.type` is used to resolve the module format of the file. - `pkg.name` is used if there is a self referencing require/import. Additionally, Node.js supports an experimental version of `Subresource Integrity` checkwhich uses the result of this package.json to validate the integrity of the file. The most important usage is that, for every `require/import` call, Node.js needs to know the module format of the file. For example, if the user require's a NPM module that uses ESM on a CommonJS (CJS) application, Node.js will need to parse the `package.json` file of that module and throw an error if the NPM package is ESM. Because of all of these calls and usages across ESM and CJS loaders, `package.json` reader is one of the most important parts of the Node.js loader implementation. ## Optimizations ### Optimizing caching layer In order to optimize the `package.json` reader performance, I first moved the caching layer to the C++ side to make the implementation be closer to the filesystem call as much as possible. This decision forced to parse the JSON file in C++. At this point, I had 2 options: - Use V8's `v8::JSON::Parse()` method which takes a `v8::String` as an input and returns a `v8::Value` as an output. - Use `simdjson` library to parse the JSON file. Since the filesystem returns a string, converting that string into a `v8::String` just to retrieve the keys and values as a `std::string` didn't make sense. Therefore, I added `simdjson` as a dependency to Node.js and used it to parse the JSON file. This change enabled us to parse the JSON file in C++ and extract and return only the necessary fields to the JavaScript side, reducing the size of the input that needs to be serialized/deserialized. ### Avoiding serialization cost In order to avoid returning unnecessary large objects, I changed the signature of the `readPackage` function to return only the necessary fields. This change simplified the `shouldUseESMLoader` as follows: ```js function shouldUseESMLoader(mainPath) { // Determine the module format of the entry point. if (mainPath && mainPath.endsWith('.mjs')) { return true } if (!mainPath || mainPath.endsWith('.cjs')) { return false } const response = getNearestParentPackageJSONType(mainPath) // No package.json or no `type` field. if (response === undefined || response[0] === 'none') { return false } const { 0: type, 1: filePath, 2: rawContent } = response checkPackageJSONIntegrity(filePath, rawContent) return type === 'module' } ``` Moving the caching layer to C++ enabled us to expose micro-functions that returns enums (integers) instead of strings to get a type of a `package.json` file. ### Reducing C++ calls to 1 to 1 On CommonJS, `readPackageConfig` is implemented on the ESM loader under `getPackageScopeConfig` function. This function made a lot of C++ calls in order to resolve and retrieve the applicable `package.json` file. The implementation was as follows: ```js function getPackageScopeConfig(resolved) { let packageJSONUrl = new URL('./package.json', resolved) while (true) { const packageJSONPath = packageJSONUrl.pathname if (packageJSONPath.endsWith('node_modules/package.json')) { break } const packageConfig = packageJsonReader.read(fileURLToPath(packageJSONUrl), { __proto__: null, specifier: resolved, isESM: true, }) if (packageConfig.exists) { return packageConfig } const lastPackageJSONUrl = packageJSONUrl packageJSONUrl = new URL('../package.json', packageJSONUrl) // Terminates at root where ../package.json equals ../../package.json // (can't just check "/package.json" for Windows support). if (packageJSONUrl.pathname === lastPackageJSONUrl.pathname) { break } } const packageJSONPath = fileURLToPath(packageJSONUrl) return { __proto__: null, pjsonPath: packageJSONPath, exists: false, main: undefined, name: undefined, type: 'none', exports: undefined, imports: undefined, } } ``` To summarize, `getPackageScopeConfig` function calls C++ 3 times from the following functions: - `new URL(...)` calls `internalBinding('url').parse()` C++ method - `path.fileURLToPath()` calls `new URL()` if the input is a string - `packageJsonReader.read()` calls `fs.readFileSync()` C++ method Moving this whole function to C++ enabled us to reduce the number of C++ calls to 1 to 1. This conversion also forced us to implement `url.fileURLToPath()` in C++. ## Results The PR that contains these changes can be found [on Github][nodejs-pr-url]. On a real-world Svelte application, the results showed 5% faster ESM execution. It also reduced the size of the cache stored by the loader by avoiding unnecessary fields. ``` ❯ hyperfine 'node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version' 'out/Release/node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version' -w 10 Benchmark 1: node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version Time (mean ± σ): 101.4 ms ± 0.6 ms [User: 96.6 ms, System: 10.8 ms] Range (min … max): 100.3 ms … 102.5 ms 28 runs Benchmark 2: out/Release/node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version Time (mean ± σ): 96.3 ms ± 0.5 ms [User: 90.9 ms, System: 10.1 ms] Range (min … max): 95.6 ms … 98.1 ms 30 runs Summary out/Release/node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version ran 1.05 ± 0.01 times faster than node ../sveltejs-realworld/node_modules/vite/dist/node/cli.js --version ``` [prior-to-optimization-run-main]: https://github.com/nodejs/node/blob/02926d3c6aaf70eba6d80423beb2d5df97e1ebc7/lib/internal/modules/run_main.js#L52 [nodejs-pr-url]: https://github.com/nodejs/node/pull/50322 [sentry-engineering-blog]: https://sentry.engineering/blog/improving-nodejs-loader-performance --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/improving-nodejs-loader-performance Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Using insecure npm package manager defaults to steal your macOS keyboard shortcuts description: "How insecure npm lifecycle-script defaults can steal macOS keyboard shortcuts, and how to lock them down." date: 2023-06-28 tag: security author: Yagiz Nizipli canonical: "https://www.yagiz.co/using-insecure-npm-defaults" markdown: "https://www.yagiz.co/using-insecure-npm-defaults.md" --- # Using insecure npm package manager defaults to steal your macOS keyboard shortcuts > How insecure npm lifecycle-script defaults can steal macOS keyboard shortcuts, and how to lock them down. *Published: 2023-06-28 · Tag: security* --- > This article is originally available at [Snyk.io][snyk-blog]. Malicious npm packages and their dangers have been a frequent topic of discussion — whether it’s [hundreds of command-and-control Cobalt Strike malware packages][cobalt-strike], [typosquatting][typosquatting], or general malware published to the npm registry (including PyPI and others). To help developers and maintainers defend against these security risks, [Snyk published a guide to npm security best practices][snyk-best-practices]. All that said, the following attack scope, which Yagiz Nizipli alerted long-time maintainers to, and the real-world risk related to data compromise are a great example of how important it is to minimize the risks of arbitrary command execution with package managers, such as those employed via npm’s postinstall lifecycle hooks. ## Life cycle scripts of npm Node Package Manager (npm) provides a set of scripts for developers and package maintainers to maintain the life cycle events of a package. These scripts provide significant value to developers by enabling them to perform various tasks or configurations as part of the package installation process. For example, with `postinstall` scripts, developers can automate tasks such as building assets, setting up environment variables, running migrations, or any tasks that can be automatically executed. The `scripts` property of a `package.json` file defines the commands triggered by the package's lifecycle and the dependent of the package you're developing. As of today, `npm` supports a limited number of life cycle scripts in any scripts property of a package.json file. For simplicity, the rest of this article will focus on the `postinstall` command. However, all concepts provided by this article also apply to other life cycle operations. ## Past security incidents There have been several high-profile incidents that had a real-world impact on JavaScript developers, including: - The [cross-env](https://iamakulov.com/notes/npm-malicious-packages/) security incident discovered by Oscar Bolmsten. - The [eslint-scope security compromise](https://eslint.org/blog/2018/07/postmortem-for-malicious-package-publishes/). - The [event-stream spear-headed attack](https://snyk.io/blog/a-post-mortem-of-the-malicious-event-stream-backdoor/) on cryptocurrency application developers. Many other JavaScript and Node.js security incidents are curated on the [Awesome Node.js Security repository][awesome-nodejs-security]. ## Data-at-rest Security Security professionals identify the protection of assets when the data is stored or at rest by `Data-at-rest`, as opposed to when it is in transit or being processed. It focuses on protecting the sensitive information stored in databases, file systems, or persistent storage. Data-at-rest security aims to prevent unauthorized access, disclosure, or data tampering while it is dormant. Various measures are available to ensure data-at-rest security, such as: - **On-demand decryption**: Decrypting only the data required to perform the current task and storing the rest of the data encrypted to prevent forbidden access. - **Access control logic**: Validating the requester's identity through a mechanism such as a password, two-factor authentication, or biometrics provided by an operating system (such as FaceID) — making it possible to limit the exposure of the resource to unwanted people. ### Attack surface of a developer Industry best practices force us to use and follow principles to develop applications. These best practices offer many advantages when working with different teams and developers but also increase the attack surface. What sort of data is lying around unencrypted in a developer machine? - Environment variables through plain text files, such as .env (available for consumption through the dotenv package). - Configuration files for projects stored as a JSON file, such as config.json. - SSH keys for accessing Github/Gitlab, which are available in the ~/.ssh folder. - And... **macOS Keyboard Shortcuts**! ## macOS Text Replacements macOS, by default, has a feature called `Text Replacements` hidden inside the system preferences applications. This feature allows users to quickly replace a word with another word. Just recently, I've learned that a developer from a well-known company was using text replacements to replace `@card` keyword with their credit card information. Even though the credit card number without the expiration date or CVV does not expose your money to outsiders, it adds an attack surface for them to exploit. ![macOS Text Replacements](https://www.yagiz.co/content/macos-text-replacements.png) > Text Replacements feature is available through System Preferences application, under the `Keyboard` menu item. ## Exfiltrating keyboard text replacements Keyboard shortcuts are stored under `defaults`, which corresponds to a filesystem backed `.plist` file somewhere in your local folder. Executing the following command will return your configured text replacements, which are also available through the System Preferences application. Remember that the following code does not require `sudo` access and can be executed by any process in your computer. ```bash title="Reading text replacements" > defaults read -g NSUserDictionaryReplacementItems ( { on = 1; replace = "@ssh-key"; with = "my-secret-password"; } ) ``` The same command can be executed through `execSync` in Node.js, and parsed without any hassle, through the `postinstall` life cycle operation supported by the `npm` package manager. The following is an example of a Node.js script that can be employed by malicious actors to access macOS text replacements and exfiltrate sensitive data: ```js title="Retrieve text replacements" import { execSync } from 'node:child_process' const decoder = new TextDecoder() const res = execSync('defaults read -g NSUserDictionaryReplacementItems') const text_replacements = decoder.decode(res) console.log(text_replacements) ``` To make sure the above code runs when this package is installed, we will update the package manifest file as follows `package.json`: ```json title="Exploiting the postinstall command" { "name": "my-useful-library", "version": "1.0.0", "description": "", "main": "index.js", "type": "module", "scripts": { "test": "echo \"Error: no test specified\" && exit 1", "postinstall": "node ./retrieve.js" }, "keywords": [], "author": "", "license": "ISC" } ``` When distributed through npm, and downloaded by a developer, this library will directly execute our custom script to retrieve and process the keyboard replacements. If you aren’t careful, it's easy to miss the line containing > node ./retrieve.js. ```bash title="NPM install flow" ➜ vulnerable npm i > my-useful-library@1.0.0 postinstall > node ./retrieve.js up to date, audited 1 package in 192ms found 0 vulnerabilities ➜ vulnerable ``` ## Protection What can you do as a developer to mitigate the security risks of malicious npm packages and general security concerns of arbitrary command execution from packages in your dependency tree? ### Ignore scripts on npm package installations Protecting yourself from packages that leverage `postinstall` scripts is possible. npm provides `--ignore-scripts` configuration when installing packages. ```bash title="Ignore all scripts" ➜ npm i --ignore-scripts up to date, audited 1 package in 124ms found 0 vulnerabilities ``` ### Use safe npm defaults NPM has a configuration file called [.npmrc](npmrc-documentation). You can change the default preferences using the `npm` CLI to ensure secure defaults: ```bash title="Ignore all scripts by default" ➜ npm config set ignore-scripts true ➜ npm i up to date, audited 1 package in 126ms found 0 vulnerabilities ``` ### Secure storage Most importantly, you should never store sensitive information in plain text. If you have to store it in plain text due to other requirements, you should always make the resource accessible through multi-factor authentication. --- **Note**: Originally, I've written this article on [Snyk blog][snyk-blog] published by Liran Tal. [awesome-nodejs-security]: https://github.com/lirantal/awesome-nodejs-security [cobalt-strike]: https://snyk.io/blog/snyk-200-malicious-npm-packages-cobalt-strike-dependency-confusion-attacks/ [typosquatting]: https://snyk.io/blog/typosquatting-attacks/ [snyk-best-practices]: https://snyk.io/blog/ten-npm-security-best-practices/ [npm-scripts-documentation]: https://docs.npmjs.com/cli/v9/using-npm/scripts [npmrc-documentation]: https://docs.npmjs.com/cli/v9/configuring-npm/npmrc [snyk-blog]: https://snyk.io/blog/using-insecure-npm-package-manager-defaults/ --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/using-insecure-npm-defaults Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: The story of WHATWG URL Specification for toddlers description: A bedtime walk through WHATWG URL edge cases — told as a story for a newborn named Ada. date: 2023-06-22 tag: experimental series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/whatwg-url-specification-for-toddlers" markdown: "https://www.yagiz.co/whatwg-url-specification-for-toddlers.md" --- # The story of WHATWG URL Specification for toddlers > A bedtime walk through WHATWG URL edge cases — told as a story for a newborn named Ada. *Published: 2023-06-22 · Tag: experimental* --- > I recently [wrote a tweet][tweet] mentioning how I tell my newborn baby, Ada, the edge cases of URL specification for her to sleep. > A good friend of mine, [James Snell, gave me the idea][james-snell-idea] to use ChatGPT to write a toddler story for URL specification. Once upon a time, in a digital world filled with wonders, there lived a bright and curious little girl named Ada. Ada loved exploring the virtual realms, where websites were like enchanted kingdoms waiting to be discovered. One day, as Ada embarked on her adventure, she stumbled upon a magical map known as the URL specification. It was a treasure trove of knowledge that revealed the secrets of navigating the Internet’s vast landscapes. Ada’s eyes sparkled with excitement as she delved into the story of URLs. She learned that these special addresses held the power to transport her to any website she desired. They were like magical keys that unlocked the doors to a world of information and imagination. As Ada delved deeper into the URL specification, she discovered the hidden components that made up these enchanting addresses. She learned about schemes, which were like magical spells that guided her web browser to the right place. There were schemes like “http://” and “https://,” each with its own purpose and adventure awaiting. Next, Ada encountered domain names, which were like unique names for the castles and kingdoms of the Internet. These names helped her find the websites she sought, such as "nodejs.org" or "yagiz.co" With a domain name, Ada could venture into realms of knowledge, creativity, and even games and stories. But the magic didn’t stop there. Ada learned about paths, which were like secret passages leading to specific pages or resources within a website. She could explore the realms of “/about” or “/blog/article-1,” uncovering hidden treasures of information and stories along the way. Ada also discovered queries, like tiny scrolls containing special instructions. These scrolls helped her communicate with websites, asking for specific things she desired. She could search for her favorite books or sort results by price using queries like “?category=books&sort=price.” With every new piece of information, Ada’s excitement grew. She uncovered fragments, like enchanted bookmarks within a webpage, guiding her to specific sections or elements of a website. She also encountered ports, which were like gateways allowing different types of communication to enter or exit a website. As Ada’s journey through the URL specification continued, she marveled at the careful craftsmanship of the Web Hypertext Application Technology Working Group (WHATWG). Their dedication and creativity had brought the URL specification to life, ensuring that everyone, including little adventurers like Ada, could navigate the Internet with ease. Armed with her newfound knowledge, Ada embarked on countless adventures, exploring websites, reading stories, and learning about the world. The URL specification became her trusted companion, empowering her to navigate the digital realm and connect with others. And so, dear Ada, with the magic of the URL specification by her side, your imagination soared, and the digital world became a boundless playground of learning and discovery. [james-snell-idea]: https://x.com/jasnell/status/1671949009538400256?s=61&t=qT2ucBMZb55QqQzOs_yctQ [tweet]: https://x.com/yagiznizipli/status/1671891859378176005?s=61&t=qT2ucBMZb55QqQzOs_yctQ --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/whatwg-url-specification-for-toddlers Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: URL specification and browser implementation differences description: "Ada, Safari, Chrome, and Firefox disagree on some URLs. Why WHATWG and browser engines drift, and what that means for parsers." date: 2023-05-25 tag: algorithms series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/url-parsing-and-browser-differences" markdown: "https://www.yagiz.co/url-parsing-and-browser-differences.md" --- # URL specification and browser implementation differences > Ada, Safari, Chrome, and Firefox disagree on some URLs. Why WHATWG and browser engines drift, and what that means for parsers. *Published: 2023-05-25 · Tag: algorithms* --- Recently, I wrote a blog post about [Ada URL parser version 2.0.0][ada-url-parser-v2]. In that post, I mentioned that the parser is fully compatible with the URL parser specification, passes all [Web Platform Tests][web-platform-tests] and currently used in [Node.js 20.0.0][node-20]. A couple of days ago, I noticed a big difference in how URLs were handled by Ada, Safari, Chrome, and Firefox. It got me thinking that it would be a good idea to explain why these differences occur and how different browsers handle URL parsing in their own unique ways, which leads to inconsistencies and compatibility issues. In this blog post, I'll break down these variations and shed light on why URLs may not always work the same way across different browsers. **Disclosure**: In the context of this blog post, I will be focusing on WHATWG URL specification, and not RFC 3986 or RFC 3987, even though historically it was written a lot earlier than the WHATWG. ## Quick Recap > Is WHATWG URL specification the only URL specification? No. One notable alternative is the URL specification maintained by the [World Wide Web Consortium][w3c]. **RFC 3986**, also known as the *Uniform Resource Identifier (URI): Generic Syntax* provides its own set of rules and guidelines for working with URLs. > Is there any differences between RFC 3986 and WHATWG? It's worth noting that the WHATWG and the W3C URL specification have some differences and have evolved separately. The WHATWG specification has been widely implemented by all web browsers, while the W3C specification is used as a reference by various web-related standards and technologies, such as cURL. Both the WHATWG and W3C specifications are important references for web developers and are widely followed in the industry. The choice of which specification to follow may depend on factors such as browser support, specific requirements of a project, or the recommendations of relevant standards organizations. > What is the meaning of WHATWG? The term **WHATWG** stands for [Web Hypertext Application Technology Working Group.][whatwg]. It is a community-driven organization that focuses on developing and maintaining web standards. The WHATWG was initially formed in response to the divergence between the World Wide Web Consortium (W3C) and the browser vendors at the time, who felt that the W3C process was too slow to address the evolving needs of web developers. > What is the URL specification? The WHATWG URL specification is a set of rules and guidelines for working with URLs (Uniform Resource Locators) in web applications. URL is a standardized way to identify and locate resources on the internet, such as web pages, images, or files. The specification provides a detailed definition of the URL syntax, parsing algorithms, and methods for manipulating URLs. The URL specification is one of the many standards maintained by the WHATWG. The group consists of a community of web developers, browser vendors, and other interested parties who collaborate to define and improve web technologies. > What is the difference between the WHATWG and the W3C? While the [W3C][w3c] is another important organization involved in web standards, the WHATWG operates independently and maintains its own set of specifications, including [the URL specification][url-spec]. ## The Problem URL standard is a living document and it gets updated quite often. Whenever a change is introduced several WHATWG members inform URL implementors about the change. Depending on the reporting method, members of WHATWG open an issue on the relevant application's bug tracker or send an email to the mailing list. Due to the priorities, some implementors may not be able to update their codebase to the latest version of the URL standard. ## Opaque hosts URLs might have protocols that does not conform to the URL standard, and in the context of specification, they are called non-special hosts. For example, HTTP, HTTPS, FTP, SSH and FILE as special protocols. We recently developed a playground for Ada URL parser, if you want to dig deeper into the result of [the URL parser][playground-opaque-host-example]. ```javascript title="Ada URL parser handling opaque-hosts" > new URL('yagiz://blog/post/1?source=rss') URL { href: 'yagiz://blog/post/1?source=rss', origin: 'null', protocol: 'yagiz:', host: 'blog', hostname: 'blog', pathname: '/post/1', search: '?source=rss', } ``` In December 28, 2016, WHATWG added the notion of opaque hosts and added support for opaque host URLs to have hostnames. This change was introduced in [Add opaque hosts pull-request][opaque-host-pr]. Unfortunately, for this particular example, browsers return different results. ### Safari Safari returns the same result as Ada URL parser, and is kept in sync with the specification. I've tested this on Safari 16.4. ### Google Chrome Returns a different hostname and pathname, and does not support opaque-hosts. [Chromium bug report][chromium-bug-report] is stil open after this change in specification was introduced back in 2016. ```javascript title="Result of Google Chrome 113.0.5672" > new URL('yagiz://blog/post/1?source=rss') URL { href: 'yagiz://blog/post/1?source=rss', origin: 'null', protocol: 'yagiz:', host: '', hostname: '', pathname: '//blog/post/1', search: '?source=rss', } ``` ### Firefox Returns the same result as Firefox, hinting towards that they're not keeping up with the specification. ```javascript title="Result of Firefox 113.0.2" > new URL('yagiz://blog/post/1?source=rss') URL { href: 'yagiz://blog/post/1?source=rss', origin: 'null', protocol: 'yagiz:', host: '', hostname: '', pathname: '//blog/post/1', search: '?source=rss', } ``` I'm afraid there are more differences than opaque-hosts and with the advancements in web3 applications, we will see more of these differences. ## Further reading I recommend reading the following blog posts by Daniel Stenberg, the author of cURL, mentioning the differences in RFC 3986 and WHATWG URL specification. - [Don't mix URL parsers][dont-mix-url-parsers] - [One URL standard please][one-url-standard-please] - [My URL isn't your URL][my-url-isnt-your-url] [ada-url-parser-v2]: https://www.yagiz.co/announcing-ada-url-parser-v2-0 [chromium-bug-report]: https://bugs.chromium.org/p/chromium/issues/detail?id=1291564&q=opaque%20host%20url&can=2 [dont-mix-url-parsers]: https://daniel.haxx.se/blog/2022/01/10/dont-mix-url-parsers/ [one-url-standard-please]: https://daniel.haxx.se/blog/2017/01/30/one-url-standard-please/ [my-url-isnt-your-url]: https://daniel.haxx.se/blog/2016/05/11/my-url-isnt-your-url/ [node-20]: https://nodejs.org/en/blog/announcements/v20-release-announce [playground-opaque-host-example]: https://playground.ada-url.com/?url=yagiz://blog/post/1?source=rss [opaque-host-pr]: https://github.com/whatwg/url/pull/185 [url-spec]: http://url.spec.whatwg.org/ [web-platform-tests]: https://web-platform-tests.org [whatwg]: https://whatwg.org/ [w3c]: https://www.w3.org/ --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/url-parsing-and-browser-differences Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Reducing the cost of string serialization in Node.js core description: "How we cut string serialization cost in Node.js URL operations, work that landed in Ada 2.0." date: 2023-04-25 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/reducing-the-cost-of-string-serialization-in-nodejs-core" markdown: "https://www.yagiz.co/reducing-the-cost-of-string-serialization-in-nodejs-core.md" --- # Reducing the cost of string serialization in Node.js core > How we cut string serialization cost in Node.js URL operations, work that landed in Ada 2.0. *Published: 2023-04-25 · Tag: performance* --- Serializing strings has been a pain point for developers, and in the context of this article, is a bottleneck in URL operations. Recently, with the help of [Daniel Lemire][daniel-lemire], we conducted an extensive research to reduce the cost of string serialization on URL parsing operations in Node.js core, resulting in a series of optimizations that addressed the issue, leading to [Ada][ada] v2.0.0. By implementing these techniques, we were able to improve the performance of URL parsing and formatting, as well as reducing memory usage and improving overall runtime stability. In this article, we will delve into the challenges we encountered while optimizing such bottlenecks in Node.js core, and try to explain the techniques we used to achieve the significant performance improvements. ## Quick Recap > What is the purpose of serialization? Serialization enables us to save the state of an object and recreate the object in a new location. In the context of this paper, serialization is required and used to pass data between C++ and JavaScript layers. > How does C++ code communicate with JavaScript code in Node.js? Node.js exposes C++ classes to the JavaScript layer using [V8][v8] through an interface called `internalBinding` where each subsystem of Node.js registers its own bindings. An example implementation of how `node:buffer` registers a certain function is available below. ```cpp title="src/node_buffer.cc" {5, 9, 12, 13} void Initialize(Local target, Local unused, Local context, void* priv) { SetMethod(context, target, "setBufferPrototype", SetBufferPrototype); } void RegisterExternalReferences(ExternalReferenceRegistry* registry) { registry->register(SetBufferPrototype); } NODE_BINDING_CONTEXT_AWARE_INTERNAL(buffer, node::Buffer::Initialize) NODE_BINDING_EXTERNAL_REFERENCE(buffer, node::Buffer::RegisterExternalReferences) ``` The internals of how `internalBinding` is created, maintained and used is out of context of this article. For more information, please refer to the Github discussion I've created called [Communication steps between JS and C++](https://github.com/orgs/nodejs/discussions/47220). ## Problem Definition Here is a quick overview of the implementation provided by Node.js v19.8.0. The code below is a simplified version of the actual implementation, and does not include base url as the parameter. Whenever a user calls `new URL` inside Node.js, the following class is created. This class is a wrapper for calling the actual implementation in C++ provided by [the Ada URL parser][ada]. The following code is available on [Github](https://github.com/anonrig/node/blob/26a967f6d0d111d72da6ea34634854e2ba8c517f/lib/internal/url.js#L567). ```js title="Node.js URL class constructor" {8} const { parse } = internalBinding('url'); class URL { #context = new URLContext(); constructor(input) { input = `${input}`; if (!parse(input, this.#onParseComplete)) { throw new ERR_INVALID_URL(input); } } } ``` The `parse` method takes 2 parameters, input and the completion callback. This is mostly done to avoid the overhead of creating a new object for each function. For example, the following code is slow due to the serialization cost of objects: ```js title="Object serialization example" {3} const parse = internalBinding('url'); const url = 'https://www.yagiz.co'; const { isValid, ...rest } = parse(url); if (isValid) { console.log(`parsed href is ${rest.href}`); } ``` In the example above, the `parse` function returns a boolean `isValid` and other properties of the parsed URL. However, the `parse` function returns these properties regardless of the `isValid` flag. This means that the structure of the `rest` object is unknown on the compile time, and V8 has to do its magic to optimize it with its limited knowledge on the executed code block. **This is a very common problem with JIT (Just in time) compilers.** Let's dive into the details of how the `parse` function is implemented: The URL constructor by default calls a C++ function called `parse` which is defined inside `src/node_url.cc`. The `parse` function is defined as follows: ```cpp title="Parse" void Parse(const FunctionCallbackInfo& args) { CHECK_GE(args.Length(), 2); CHECK(args[0]->IsString()); // input CHECK(args[1]->IsFunction()); // complete callback Local success_callback_ = args[2].As(); Environment* env = Environment::GetCurrent(args); HandleScope handle_scope(env->isolate()); Context::Scope context_scope(env->context()); Utf8Value input(env->isolate(), args[0]); ada::result out = ada::parse(input.ToStringView()); if (!out) { return args.GetReturnValue().Set(false); } auto argv = GetCallbackArgs(env, out); USE(success_callback_->Call( env->context(), args.This(), argv.size(), argv.data())); args.GetReturnValue().Set(true); } ``` Whenever the parse function is called, it needs to be called with `input` parameter which is a string, and a callback function to pass the values back to the JavaScript layer. This is a smart way of solving the serialization problem of objects, and it is also a very common pattern in Node.js core. Unfortunately, this pattern leads to making this function a function that has a side effect. Meaning, it has to mutate the callback according to the result of the parsing. Here's an example of how the callback is mutated to return the result to the JavaScript layer: ```cpp title="GetCallbackArgs" {8-19} auto GetCallbackArgs(Environment* env, const ada::result& url) { Local context = env->context(); Isolate* isolate = env->isolate(); auto js_string = [&](std::string_view sv) { return ToV8Value(context, sv, isolate).ToLocalChecked(); }; return std::array{ js_string(url->get_href()), js_string(url->get_origin()), js_string(url->get_protocol()), js_string(url->get_hostname()), js_string(url->get_pathname()), js_string(url->get_search()), js_string(url->get_username()), js_string(url->get_password()), js_string(url->get_port()), js_string(url->get_hash()), }; } ``` In order to process and save this data on the JavaScript layer, preferably in URL class, JavaScript layer had to have a `heavy` function to update the current context of the URL: ```js title="Simplified version of the URL class onParseComplete function" {1-2} #onParseComplete = (href, origin, protocol, hostname, pathname, search, username, password, port, hash) => { this.#context.href = href; this.#context.origin = origin; this.#context.protocol = protocol; this.#context.hostname = hostname; this.#context.pathname = pathname; this.#context.search = search; this.#context.username = username; this.#context.password = password; this.#context.port = port; this.#context.hash = hash; }; ``` This implementation as you've realized is not very efficient. It requires sharing the knowledge of the callback function parameters, by index, between JavaScript and C++. On top of this being a bad practice, there is a lot of room for improvement in terms of performance. The bridge between C++ to JavaScript is not very efficient, leading to bottlenecks when used in hot paths. This wasn't a problem until now, where the true performance cost of this function lied in the fact that the URL parser was slow. However, with [Ada URL parser][ada] the bottleneck was moved to the serialization of the result. As you know the URL contains a lot of properties, where `href` is the only attribute that contains all of the properties of URL, hence the identifier of the URL. ```shell title="URL properties" > new URL('https://www.yagiz.co') URL { href: 'https://www.yagiz.co/', origin: 'https://www.yagiz.co', protocol: 'https:', username: '', password: '', host: 'www.yagiz.co', hostname: 'www.yagiz.co', port: '', pathname: '/', search: '', searchParams: URLSearchParams {}, hash: '' } ``` As you might notice, origin, protocol, host, hostname and others are all substrings of `href`. Well, the solution is not as simple as this, because the `origin` might differ from `URL` where the `hostname` can be different with different `pathname` values. There are lots of edge cases that needs to be resolved if we are going to resolve this. ## The Solution With [Ada URL Parser v2.0.0][ada], we incorporated a common approach in industry for storing the URL properties. The idea is to store the href, and use offsets to represent the URL properties. This way, we can have access to the URL properties without knowing the **business logic** behind "How to parse a URL?". This solution comes with another advantage on top of solving the serialization cost. The parsing becomes faster, because we don't need to create multiple strings for each URL property. We can reserve and allocate a string with a guessed size, and use the offsets to construct the `href` while parsing the URL. This reduces the memory allocations, and the time spent on parsing the URL. ### URL Components Here's a quick recap from [Ada v2.0 article][ada-v2-article]: ```plaintext title="URL Components structure" https://user:pass@example.com:1234/foo/bar?baz#quux | | | | ^^^^| | | | | | | | | | `----- hash_start | | | | | | `--------- search_start | | | | | `----------------- pathname_start | | | | `--------------------- port | | | `----------------------- host_end | | `---------------------------------- host_start | `--------------------------------------- username_end `--------------------------------------------- protocol_end ``` The structure of the URL class stayed the same, but with little caveats. On the C++ side, we created a class called `BindingData`. ```cpp title="src/node_url.h" class BindingData : public SnapshotableObject { public: // This is a simplified version of the class static void Parse(const v8::FunctionCallbackInfo& args); static void Initialize(v8::Local target, v8::Local unused, v8::Local context, void* priv); static void RegisterExternalReferences(ExternalReferenceRegistry* registry); private: static constexpr size_t kURLComponentsLength = 9; AliasedUint32Array url_components_buffer_; }; ``` ### `Bindingdata` `Bindingdata` class is initialized and snapshotted in the build time and is used to store an `AliasedUint32Array` called `url_components_buffer_` with a length of `9` unsigned integers. This property will be used to store the offsets of the URL. Due to the single-threaded environment and the non-parallel execution of the URL parser, we ensure that that only **1** Uint32Array is created for parsing URLs throughout the lifecycle of the Node.js application. > What is AliasedUint32Array? Referencing from the implementation itself: `AliasedUint32Array` is a class that encapsulates the technique of having a native buffer mapped to a JavaScript object. Writes to the native buffer can happen efficiently without going through JavaScript, and the data is then available to user via the exposed JavaScript object. While this technique is computationally efficient, it is effectively a write to JavaScript application state without going through the monitored API. Thus any VM capabilities to detect the modification are circumvented. The implementation is available at [Github][nodejs-aliased-buffer]. Here's the implementation of the parse function from `BindingData` class. ```cpp title="src/node_url.cc" void BindingData::Parse(const FunctionCallbackInfo& args) { CHECK_GE(args.Length(), 1); CHECK(args[0]->IsString()); // input // args[1] // base url BindingData* binding_data = Realm::GetBindingData(args); Environment* env = Environment::GetCurrent(args); HandleScope handle_scope(env->isolate()); Context::Scope context_scope(env->context()); Utf8Value input(env->isolate(), args[0]); auto out = ada::parse(input.ToStringView()); if (!out) { return args.GetReturnValue().Set(false); } binding_data->UpdateComponents(out->get_components(), out->type); args.GetReturnValue().Set( ToV8Value(env->context(), out->get_href(), env->isolate()) .ToLocalChecked()); } ``` As a result of this optimization, we just need to update the `url_components_buffer_` with the offsets of the URL properties. This is done by the `binding_data->UpdateComponents` method. ### JavaScript Class Let's dive into the JavaScript implementation of the URL class. Here's the implementation as of Node.js 20 (April 24, 2023). ```js title="URL class" const bindingUrl = internalBinding('url'); class URL { #context = new URLContext(); constructor(input, base = undefined) { input = `${input}`; const href = bindingUrl.parse(input, base); if (!href) { throw new ERR_INVALID_URL(input); } this.#updateContext(href); } } ``` The `parse` method returned a string or an undefined value depending on the success of the parsing function. The string value represents the `href` part of the URL. Immediately after successful parsing, `this.#updateContext(href)` method is called to access the URL components (indexes of the URL properties) and update the current url instance context. As a result, the cost of parsing an invalid URL has significantly decreased. ### Updating the URL context The following code is triggered every time, a URL setter is triggered, as well as the everytime a URL is constructed. ```js title="Simplified version of #updateContext(href) implementation" /bindingUrl.urlComponents/ #updateContext(href) { const { 0: protocol_end, 1: username_end, 2: host_start, 3: host_end, 4: port, 5: pathname_start, 6: search_start, 7: hash_start, 8: scheme_type, } = bindingUrl.urlComponents; this.#context.protocol_end = protocol_end; this.#context.username_end = username_end; this.#context.host_start = host_start; this.#context.host_end = host_end; this.#context.port = port; this.#context.pathname_start = pathname_start; this.#context.search_start = search_start; this.#context.hash_start = hash_start; this.#context.scheme_type = scheme_type; } ``` Due to the object destructure of `bindingUrl.urlComponents`, the barrier between JavaScript and C++ is only crossed once, reducing the performance cost of string serialization. The usage of indexes as offsets to access the URL properties through a string is not a new idea. As most of you know, it's called lazy loading. > What is lazy loading? In the context of algorithms, lazy loading is a strategy for optimizing performance by deferring the calculation of values until they are actually needed. This is often used in cases where computing all possible values in advance would be inefficient or impractical. Instead, the algorithm only calculates values as they are requested, often caching the results for future use. This approach can help reduce the amount of computation required, improve memory usage, and speed up the overall execution time of the algorithm. The usage of lazy-loading forces us to know the context of where lazy loading is used. In the case of the URL class, the cost of parsing vs. the cost of accessing the URL properties is the main factor that we need to consider. As a result of this experiment, and optimizations done on both [Ada][ada] and Node.js, we were able to reduce the performance cost of parsing by a significant amount. Here's the result of parsing 100,000 URLs in different Node.js versions on M1 Pro Max. The benchmark code is available at [Github][node-benchmarks]. | Runtime | Ada Version | Time (ms/iter) | | --------------- | ----------- | -------------- | | Node.js 20 | 2.0.0 | 38.97 | | Node.js 19.8.1 | 1.0.4 | 79.29 | | Node.js 19.7.0 | 1.0.1 | 105.59 | | Node.js 19.6.1 | - | 140.42 | If you have a passion for performance and Node.js, we are actively looking for contributors for our [performance team][nodejs-performance-team]. [ada]: https://github.com/ada-url/ada [ada-v2-article]: https://www.yagiz.co/announcing-ada-url-parser-v2-0 [nodejs-aliased-buffer]: https://github.com/nodejs/node/blob/main/src/aliased_buffer.h#L32 [daniel-lemire]: https://lemire.me [node-benchmarks]: https://github.com/anonrig/node-benchmarks [nodejs-performance-team]: https://github.com/nodejs/performance [v8]: https://v8.dev --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/reducing-the-cost-of-string-serialization-in-nodejs-core Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Securing your Next.js 13 application description: "Security headers, CSP, and CSRF protections that matter in a Next.js 13 app." date: 2023-04-22 tag: security author: Yagiz Nizipli canonical: "https://www.yagiz.co/securing-your-nextjs-13-application" markdown: "https://www.yagiz.co/securing-your-nextjs-13-application.md" --- # Securing your Next.js 13 application > Security headers, CSP, and CSRF protections that matter in a Next.js 13 app. *Published: 2023-04-22 · Tag: security* --- Next.js is a popular framework for building server-side rendered React applications, but like any web application, it's crucial to take security seriously. In this blog post, we will discuss some essential security practices that you can implement in your Next.js application. We will cover topics such as security headers, including Content Security Policy, Referrer-Policy, and X-Frame-Options, as well as preventing cross-site request forgery attacks. By following these security best practices, you can help ensure the safety and privacy of your users' sensitive information. ## Security Headers One of the essential security practices for Next.js applications is to implement recommended security headers. Security headers are HTTP response headers that allow web developers to control and enforce additional security mechanisms in their applications. The code snippet below shows an example of some recommended security headers to add to your Next.js application. ```js title="Adding security headers to your next.config.mjs file" const ContentSecurityPolicy = ` default-src 'self' vercel.live; script-src 'self' 'unsafe-eval' 'unsafe-inline' cdn.vercel-insights.com vercel.live; style-src 'self' 'unsafe-inline'; img-src * blob: data:; media-src 'none'; connect-src *; font-src 'self'; `.replace(/\n/g, ''); const securityHeaders = [ { key: 'Content-Security-Policy', value: ContentSecurityPolicy }, { key: 'Referrer-Policy', value: 'origin-when-cross-origin' }, { key: 'X-Frame-Options', value: 'DENY' }, { key: 'X-Content-Type-Options', value: 'nosniff' }, { key: 'X-DNS-Prefetch-Control', value: 'on' }, { key: 'Strict-Transport-Security', value: 'max-age=31536000; includeSubDomains; preload' }, { key: 'Permissions-Policy', value: 'camera=(), microphone=(), geolocation=()' }, ]; /** @type {import('next').NextConfig} */ export default { headers() { return [ { source: '/(.*)', headers: securityHeaders }, ]; } } ``` ### Content Security Policy Content Security Policy (CSP) is a security header that allows you to specify a set of rules that define the sources from which the application can load resources such as scripts, styles, and images. It is designed to mitigate cross-site scripting (XSS) attacks by preventing the execution of scripts from untrusted sources. The value of the Content-Security-Policy header in the code snippet above specifies the allowed sources for various resources. ### Referrer-Policy Referrer-Policy is a security header that allows you to control how much information is passed in the HTTP Referer header when navigating from one page to another. The value of the Referrer-Policy header in the code snippet above specifies that the origin should be sent as the referrer when the request is made from the same origin, but not when the request is made from a different origin. ### X-Frame-Options X-Frame-Options is a security header that prevents clickjacking attacks by ensuring that a webpage can only be displayed in a frame or iframe if it is from the same origin. The value of the X-Frame-Options header in the code snippet above specifies that the page should not be displayed in any frame or iframe. ### X-Content-Type-Options X-Content-Type-Options is a security header that prevents MIME type sniffing, a vulnerability that allows attackers to trick the browser into interpreting a resource as a different MIME type, possibly executing malicious code. The value of the X-Content-Type-Options header in the code snippet above specifies that the browser should not try to guess the MIME type and should only use the declared MIME type. ### X-DNS-Prefetch-Control X-DNS-Prefetch-Control is a security header that controls whether the browser should prefetch DNS entries for links on a webpage. Prefetching DNS entries can reduce the latency of subsequent requests but can also be used for tracking purposes. The value of the X-DNS-Prefetch-Control header in the code snippet above specifies that DNS prefetching should be enabled. ### Strict-Transport-Security Strict-Transport-Security (HSTS) is a security header that enforces the use of HTTPS by the browser and prevents downgrade attacks. The value of the Strict-Transport-Security header in the code snippet above specifies that the browser should use HTTPS for all requests for the next 365 days, including subdomains, and preload the HSTS policy to all browsers. ### Permissions-Policy Permissions-Policy is a security header that allows you to specify which APIs and features are allowed to be used in the application. The value of the Permissions-Policy header in the code snippet above specifies that the camera, microphone, and geolocation APIs are not allowed to be used. ## Cross Site Request Forgery **Cross-site request forgery (CSRF)** is a type of attack that exploits the trust between a user and a website to perform unauthorized actions on the user's behalf. CSRF attacks can lead to serious security breaches, such as account takeover, data theft, and malware installation. In this type of attack, an attacker tricks a user into executing a malicious action on a website without their knowledge or consent. Since the attack is carried out with the user's credentials, it can be difficult to detect and mitigate. ### Implement on your own Here is a high level implementation 1. Create a `middleware.ts` on your root folder. 2. Filter middleware to grab only `GET` requests 3. Create a secure cookie with a short expiration date with a value that is signed by the secret key only known by the server-side available environment key. 4. Update the NextResponse headers with `X-CSRF-Token` header with the raw value 5. On the server-side of the page you want to add CSRF protection, grab the http value using `headers().get('X-CSRF-Token')` and pass it to the client-side component. 6. On the client-side of the page, create a hidden input with the value of the `X-CSRF-Token` header. 6. On your **route endpoint**: 1. Validate that the CSRF header exists 2. Validate that a cookie is sent with the existing request 3. Sign `X-CSRF-Token` value with your secret key and make sure that it is same as the cookie of the request. 4. If any of these validations fail, throw an error. ### Use a package 1. Install [`edge-csrf`](https://github.com/amorey/edge-csrf) module using `npm install edge-csrf` or `yarn add edge-csrf`. 2. Create a `middleware.ts` on your root folder. ```js title="middleware.ts" import csrf from 'edge-csrf'; import { NextResponse } from 'next/server'; import type { NextRequest } from 'next/server'; // initalize protection function const csrfProtect = csrf({ cookie: { secure: process.env.NODE_ENV === 'production', }, }); export async function middleware(request: NextRequest) { const response = NextResponse.next(); // csrf protection const csrfError = await csrfProtect(request, response); // check result if (csrfError) { return new NextResponse('invalid csrf token', { status: 403 }); } return response; } ``` 3. And on your React component: ```js title="pages/page.ts" import type { NextPage, GetServerSideProps } from 'next'; import React from 'react'; type Props = { csrfToken: string; }; export const getServerSideProps: GetServerSideProps = async ({ res }) => { const csrfToken = res.getHeader('x-csrf-token') || 'missing'; return { props: { csrfToken } }; } const FormPage: NextPage = ({ csrfToken }) => { return (
); } export default FormPage; ``` 4. Create your API endpoint: ```js title="pages/api/form-handler.ts" import type { NextApiRequest, NextApiResponse } from 'next'; type Data = { status: string }; export default function handler(req: NextApiRequest, res: NextApiResponse) { // this code won't execute unless CSRF token passes validation res.status(200).json({ status: 'success' }); } ``` --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/securing-your-nextjs-13-application Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Announcing Ada URL parser v2.0 description: "Ada 2.0 drops the ICU dependency, adds url_aggregator for one-shot parses, and roughly doubles throughput on some workloads." date: 2023-03-30 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/announcing-ada-url-parser-v2-0" markdown: "https://www.yagiz.co/announcing-ada-url-parser-v2-0.md" --- # Announcing Ada URL parser v2.0 > Ada 2.0 drops the ICU dependency, adds url_aggregator for one-shot parses, and roughly doubles throughput on some workloads. *Published: 2023-03-30 · Tag: performance* --- [Ada URL Parser](https://github.com/ada-url/ada), a powerful tool for parsing URLs, has just been updated to version 2.0 after the release of version 1.0.4 just a month ago. This latest version brings some significant improvements over its predecessor, includinga doubling of execution speed in some cases, as well as reduced memory usage and allocations. These enhancements make the Ada URL Parser more efficient and capable of handling a broader range of URL parsing tasks with ease. In this blog post, we'll take a closer look at the new features and performance improvements of the Ada URL Parser v2.0 and explore how they can benefit developers in their everyday work. In addition to the performance and memory improvements, the Ada URL Parser v2.0 also introduces a new feature that will be of particular interest to developers working with one-time URL parsing tasks. With this update, Ada now supports two different implementations of URL parsing out of the box. ## Dropping ICU requirement [The International Components for Unicode (ICU) library](https://icu.unicode.org/) was a vital component of the Ada URL Parser v1.x and was required for proper URL **hostname** parsing and manipulation. This is because URLs can contain a wide variety of international characters, including non-Latin alphabets, and the ICU library provides a comprehensive set of tools for handling these characters in a standardized manner. Specifically, the ICU library provides support for Unicode normalization, which is a critical step in URL parsing. Normalization ensures that URLs containing international characters are properly encoded and can be processed correctly by other software components. Additionally, the ICU library provides support for character set conversion and detection, which are essential for handling URLs from different regions of the world. Prior to v2.0, Ada was dependent on ICU for `to_ascii` and `to_unicode` operations. The ICU library is available for all systems and we cannot guaranteed that it is up-to-date. Providing our own unicode functions allows us to fully support the standard across a broad range of systems. Furthermore, our benchmarks reveal that we sometimes achieve better performance with our own dedicated functions. With Ada v2.0, we are dropping the system requirement of ICU and rolling out our implementation of the Unicode Specification, and releasing it publicly on GitHub ([ada/idna](https://github.com/ada-url/idna)). ## Introducing url\_aggregator The first implementation, **ada::url**, is the existing URL representation that was suitable and optimized for environments where the reference to the instance can persist and live on. It provides a comprehensive set of URL parsing functions and is suitable for tasks that involve ongoing URL manipulation. ```cpp title="ada::url example" auto url = ada::parse("https://www.google.com"); url->set_pathname("/my-super-long-path") // url->get_pathname() will return "/my-super-long-path" ``` The second implementation, **ada::url\_aggregator**, is a new addition to the Ada URL Parser family. It is explicitly designed for parsing and deserializing URLs in one-time environments, where a new URL instance is created for each parsing operation. This implementation uses a similar URL representation inspired by Servo URL parser making it an ideal choice for performance-critical applications. ```cpp title="ada::url_aggregator example" auto url = ada::parse("https://www.google.com"); url->set_pathname("/my-super-long-path") // url->get_pathname() will return "/my-super-long-path" ``` As you may have realized, the public API of **ada::url** and **ada::url\_aggregator** is same. However, their internal implementations and the way in which each URL subcomponent is defined is the key differentiator. The **ada::url** structure stores the components of the parsed URL in different string instances, making updates fast. The **ada::url\_aggregator** uses a single string buffer, thus minimizing memory usage, at the expense of more work during updates ### Reducing string allocations in Ada One of the key features of the **Ada URL Parser v2.0** is its ability to provide a comprehensive representation of the various components of a URL string. This representation is based on the [WHATWG URL specification](https://url.spec.whatwg.org/) and includes a range of offsets for different URL components. These offsets are used to identify the start and end indexes of different parts of the URL, such as the protocol, username, hostname, port, pathname, search, and hash. By using these offsets, developers can easily extract and manipulate different parts of a URL string without having to worry about the underlying parsing details. ```plaintext title="URL Components structure" https://user:pass@example.com:1234/foo/bar?baz#quux | | | | ^^^^| | | | | | | | | | `----- hash_start | | | | | | `--------- search_start | | | | | `----------------- pathname_start | | | | `--------------------- port | | | `----------------------- host_end | | `---------------------------------- host_start | `--------------------------------------- username_end `--------------------------------------------- protocol_end ``` For example, the **ada::url\_components** feature includes the protocol\_end offset, which represents the ending index of the protocol component in the URL string. It also includes the username\_end offset, which is used for URLs that contain a username. Additionally, the feature provides host\_start and host\_end offsets, which represent the start and end indexes of the hostname component of the URL. The port, pathname\_start, search\_start, and hash\_start offsets are also included, allowing developers to easily extract and manipulate the corresponding URL components. ## Benchmarks We have been actively working on improving the library's benchmark infrastructure to provide a more realistic comparison of Ada's performance compared to other similarly-scoped libraries. As part of this effort, we developed a crawler that visited 100,000 URLs from the most visited 100 websites, with a limit of 100 URLs per unique domain. The crawler was designed to simulate real-world URL parsing and manipulation scenarios and provided valuable insights into the performance of the Ada URL Parser in comparison to other libraries. All benchmarks are executed using the Apple M1 Max processor. ### Comparing Ada with alternatives The following [benchmark code](https://github.com/ada-url/ada/tree/main/benchmarks) is available on GitHub as well as [the dataset](https://github.com/ada-url/url-dataset) of 100,000 URLs. We used cURL 8.0.1, Servo 2.3.1 (using Rust 1.64.0), and Boost 1.81.0 for benchmarking. ```plaintext title="Benchmark results of Ada" Benchmark time/url url/s ------------------------------------------------- ada::url_aggregator 222.298ns 4.49846M/s ada::url 283.211ns 3.53093M/s Boost 335.577ns 2.97994M/s Servo 686.495ns 1.45667M/s cURL 1.32924us 752.31k/s ``` Ada is currently **50% faster than Boost**, **3x faster than Servo** and **6x faster than cURL** in the given dataset. **_Disclaimer:_**cURL _follows [RFC 3986+](https://curl.se/docs/url-syntax.html), and Boost follows [RFC 3986](https://www.rfc-editor.org/rfc/rfc3986)._ ### Comparing Node.js (with Ada) with alternatives The following benchmark code is available on [GitHub](https://github.com/anonrig/node-benchmarks/tree/main/url) as well as [the dataset](https://github.com/ada-url/url-dataset) of 100,000 URLs. We used Node.js main branch, Bun 0.5.8 and Deno 1.32.1 for benchmarking. ```plaintext title="Comparing Node.js with Bun and Deno" benchmark time (avg) (min … max) ------------------------------------------------------------ Node 41.08 ms/iter (40.97 ms … 41.32 ms) Bun 75.31 ms/iter (73.73 ms … 80.3 ms) Deno 118.66 ms/iter (118.41 ms … 118.92 ms) ``` Node.js is currently **82% faster than Bun** and **3x faster than Deno** in the given dataset. All this great work is a result of a collaboration between [Daniel Lemire](https://github.com/lemire), [Miguel Teixeira](https://github.com/miguelteixeiraa), and me ([Yagiz Nizipli](https://github.com/anonrig)). The full changelog can be found on [Ada's GitHub releases page](https://github.com/ada-url/ada/releases). --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/announcing-ada-url-parser-v2-0 Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Performance tips for C++ developers description: "Practical C++ performance notes: avoid copies, pick the right container, and let the compiler see the hot path." date: 2023-03-24 tag: performance author: Yagiz Nizipli canonical: "https://www.yagiz.co/performance-tips-for-c-developers" markdown: "https://www.yagiz.co/performance-tips-for-c-developers.md" --- # Performance tips for C++ developers > Practical C++ performance notes: avoid copies, pick the right container, and let the compiler see the hot path. *Published: 2023-03-24 · Tag: performance* --- C++ is a powerful and versatile programming language that is widely used for a variety of applications, including system programming, game development, and scientific computing. However, writing efficient and performant C++ code can be a challenging task, especially when dealing with large and complex projects. In this blog post, we will explore all possible performance optimization techniques for C++ developers, from algorithmic optimizations to low-level code optimizations. We will cover a wide range of topics, including memory management, data structures, parallel processing, and compiler optimizations. Whether you are a seasoned C++ developer or just getting started, this post will provide you with a comprehensive guide to optimizing your C++ code for maximum performance. ## Algorithms **Use efficient algorithms and data structures** Make sure you're using the most efficient algorithms and data structures for the problem you're solving. For example, if you need to search through a large collection of data frequently, consider using a hash table or a balanced tree instead of a linear search. ## Avoid unnecessary copies C++ is a pass-by-value language, so when you pass an object to a function, a copy is made. This can be expensive, especially for large objects. To avoid unnecessary copies, use references or pointers instead. ```cpp void process_string(const std::string& str) { // Do something with the string } int main() { std::string my_string = "Hello, world!"; process_string(my_string); // Pass the string by reference return 0; } ``` In this example, the `process_string` function takes a `const` reference to a `std::string` parameter instead of taking the `std::string` object itself. This avoids making a copy of the `std::string` object when it's passed to the function. If we had defined `process_string` like this instead: ```cpp void process_string(std::string str) { // Do something with the string } ``` Then calling `process_string(my_string)` would create a copy of `my_string`, which can be expensive if `my_string` is large. By using a reference parameter instead, we avoid the copy and improve performance. ## Usage of const and constexpr 1. Use **const** whenever possible to prevent accidental modification of variables and improve the compiler's ability to optimize your code. 2. Use **constexpr** for values that can be computed at compile-time, which can help the compiler optimize your code further. ## Avoid virtual functions When a virtual function is called on an object, the C++ runtime needs to look up the correct function implementation in a virtual function table, or vtable, which is a data structure that stores pointers to the virtual functions for a class. This lookup involves following a pointer to the vtable for the object's class, and then following another pointer in the vtable to the correct function implementation. This extra indirection can cause a small performance overhead compared to non-virtual function calls, which can be called directly without any vtable lookup. Additionally, virtual function calls can't be easily inlined by the compiler, since the function to be called may not be known until runtime. Inlining a function can often provide performance benefits by reducing the overhead of the function call itself, but this isn't possible with virtual functions. That said, the performance overhead of virtual functions is often small, and they're a powerful tool for implementing polymorphism and dynamic dispatch in C++. In many cases, the benefits of virtual functions outweigh the performance costs, and it's generally best to use virtual functions when they're appropriate for your design. If possible, use templates or function overloading instead. ## Use optimizations flags Most compilers have optimization flags that can significantly improve performance. For example, use -O3 with GCC or Clang to enable the highest level of optimization. GCC provides a wide range of optimization flags that can help you optimize your C++ code. Here are some of the most commonly used optimization flags: 1. \-O: Enables basic optimization that can improve code size and execution time. 2. \-O1, -O2, -O3: Enables progressively more aggressive optimization levels that can improve code performance, but may also increase compilation time and code size. 3. \-Os: Optimizes for code size by performing aggressive code optimizations that reduce the size of the executable. 4. \-Ofast: Enables aggressive optimization that can improve code performance, but may also sacrifice accuracy and correctness in certain cases. 5. \-march: Specifies the target processor architecture for code generation, allowing the compiler to generate code that takes advantage of specific processor features. 6. \-mtune: Specifies the target processor model for code generation, allowing the compiler to tune the code for a specific processor model. 7. \-funroll-loops: Enables loop unrolling, which can improve performance by reducing the overhead of loop iterations. 8. \-finline-functions: Enables function inlining, which can improve performance by reducing the overhead of function calls. 9. \-fprofile-generate/-fprofile-use: Enables profile-guided optimization (PGO), which uses information from a previous profiling run to optimize the code for improved performance. 10. \-flto: Enables link-time optimization (LTO), which performs optimization across multiple translation units, allowing the compiler to perform more aggressive optimizations. These are just some of the many optimization flags available on GCC. The specific flags you use will depend on the needs of your application and the target platform you're optimizing for. ## Profile your code Use a profiler to identify performance bottlenecks in your code. This will help you focus your optimization efforts on the parts of your code that will have the biggest impact. There are several profiling tools available for profiling C++ code. Here are some popular ones: 1. Valgrind: Valgrind is a powerful profiling tool that can detect memory leaks, thread errors, and other performance issues. It provides a suite of tools, including Memcheck, which can help you identify memory errors in your C++ code. 2. Google Performance Tools (gperftools): gperftools is a collection of profiling and performance analysis tools for C++ code. It includes a CPU profiler, a heap profiler, and a heap-checker tool that can help you identify memory leaks and other memory errors. 3. Intel VTune: Intel VTune is a performance analysis tool that can help you analyze and optimize the performance of C++ code running on Intel processors. It provides a wide range of profiling features, including CPU profiling, memory profiling, and thread profiling. 4. Linux perf: Linux perf is a performance analysis tool that is included in the Linux kernel. It provides a suite of profiling tools, including CPU profiling, memory profiling, and kernel tracing, that can help you identify performance bottlenecks in your C++ code. ## Optimize for target platform Make sure you're optimizing your code for the platform it will run on. This includes using platform-specific optimizations and avoiding platform-specific inefficiencies. To optimize your C++ code for target platforms, you can follow these steps: 1. **Understand the target platform:** To optimize for a specific platform, you need to understand its hardware architecture, instruction set, memory hierarchy, and other performance characteristics. This can help you identify performance bottlenecks and opportunities for optimization. 2. **Use platform-specific optimization flags:** Most compilers support platform-specific optimization flags that can help you optimize your code for a specific platform. For example, GCC supports the `-march` and `-mtune` flags, which allow you to specify the target processor architecture and tune the code for a specific processor model, respectively. 3. **Use SIMD instructions:** Single Instruction Multiple Data (SIMD) instructions can help you perform the same operation on multiple data elements simultaneously, which can significantly improve performance for certain types of computations. Most modern processors support SIMD instructions through instruction sets such as SSE, AVX, or NEON. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/performance-tips-for-c-developers Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Working with 1000+ tests on a stable C++ library description: How we refactored Ada's C++ URL parser under a 1000+ test suite without breaking WHATWG edge cases. date: 2023-03-18 tag: coding series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/working-with-1000-tests" markdown: "https://www.yagiz.co/working-with-1000-tests.md" --- # Working with 1000+ tests on a stable C++ library > How we refactored Ada's C++ URL parser under a 1000+ test suite without breaking WHATWG edge cases. *Published: 2023-03-18 · Tag: coding* --- Working on a project that handles every edge case of a URL specification and has over 1000 unit tests can be quite challenging. Recently, I had the opportunity to work on a refactor and a new API for [the Ada URL parser library](https://github.com/ada-url/ada) with Daniel Lemire. This project presented interesting challenges that tested the limits of our motivation, determination, and productivity. In this blog post, I will share my experience working on a stable and well-tested C++ project and discuss some of the issues we faced along the way. ## Story begins Our overall test suite for Ada, consists of 5+ files, running using CTest. For reasons which are unrelevant of this article, we've implemented our own macros for assertion, succeeding and failing the test suites. Here's an exaple of a C++ macro for our assertion function. ```cpp title="Definition of the TEST_ASSERT macro" #define TEST_ASSERT(LHS, RHS, MESSAGE) \ do { \ if (LHS != RHS) { \ std::cerr << "Mismatch: '" << LHS << "' - '" << RHS << "'" << std::endl; \ TEST_FAIL(MESSAGE); \ } \ } while (0); ``` Knowing the internals of how our tests are executed helped a new C++ developer like me and increased the onboarding process. ## Logging As we progressed through the development process, we found that we were spending a lot of time on debugging, so we made the decision to implement our own logging system to capture valuable information. With this approach, whenever we encountered a bug, we could easily run the test suite and debug values through the terminal output. Daniel played a crucial role in this effort and was able to add a simple yet effective logger to the project in no time at all! ```cpp title="Definition of our internal log function" template ada_really_inline void log([[maybe_unused]] T t) { #if ADA_LOGGING std::cout << "ADA_LOG: " << t << std::endl; #endif ``` At the time, we were in the process of implementing the first version of Ada, so it wasn't an issue that the logs prior to the failed test were irrelevant. In fact, this helped us to pinpoint the root cause of the issue more easily. ## The struggle After releasing version 1.0 of the Ada URL parser library, we turned our attention to performance optimizations and identified a bottleneck in the string creation process. We began working on a new API that would run parallel to the existing implementation, as each is suited for different use cases. However, this decision presented a new set of challenges that we were not prepared for. Despite having started work on the new API over a month ago, we found that our productivity was decreasing and we were losing motivation. Upon discussing the matter with Daniel, we identified several key issues that were holding us back: ### Debugging **Problem:** After enabling debug mode for logging, the test runner output made it really difficult to find the starting point of a test in the logger. This could be easily resolved by a visual element, but due to the nature of our implementation, the the parser could have been called multiple times in a single test. For example; having a base URL with a input would call the URL parser 2 times. **Proposed solution:** We should only print logs when necessary, and the test runner should know the context of the log and print them to the console only when a test fails. ### Motivation **Problem:** We didn't have visibility into our progress. We only knew if we passed or failed the test suite, but didn't know how many tests were passing or how far we were from achieving our goals. **Proposed solution:** Presenting our current state to the developer using basic gamification techniques to create a dopamine effect for manipulating how we interpret our perspective towards the result. An example of this could be a progress bar, a percentage of succeeded tests, showing the difference between previous test executions with green/red highlights. ### Test Execution **Problem:** We were stopping the test suite as soon as a test failed, which prevented us from seeing the output of tests that might have been fixed by recent changes. **Proposed solution:** We should continue running the test suite even after encountering a failed test, so that we can see the output of all tests and ensure that we haven't introduced any new errors. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/working-with-1000-tests Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Using V8 Fast API in Node.js core description: "Embedder functions implemented in C++ incur a high overhead, so V8 provides an API to implement fast-path C functions which may be invoked directly from JITted code." date: 2023-01-13 tag: coding author: Yagiz Nizipli canonical: "https://www.yagiz.co/using-v8-fast-api-in-node-js-core" markdown: "https://www.yagiz.co/using-v8-fast-api-in-node-js-core.md" --- # Using V8 Fast API in Node.js core > Embedder functions implemented in C++ incur a high overhead, so V8 provides an API to implement fast-path C functions which may be invoked directly from JITted code. *Published: 2023-01-13 · Tag: coding* --- Node.js uses [V8](https://github.com/v8/v8) as the JavaScript engine. Embedder functions implemented in C++ incur a high overhead, so V8 provides an API to implement fast-path C functions which may be invoked directly from JITted code. These functions also come with additional constraints, for example, they may not trigger garbage collection. ### Limitations * Fast API functions may not trigger garbage collection. This means by proxy that JavaScript execution and heap allocation are also forbidden, including `v8::Array::Get()` or `v8::Number::New()` * Throwing errors is not available on fast API but can be done through the fallback to slow API. * Not all parameter and return types are supported in fast API calls. For a full list, please look into [v8-fast-api-calls.h](https://source.chromium.org/chromium/chromium/src/+/main:v8/include/v8-fast-api-calls.h). ### Requirements * Each unique fast API function signature should be defined inside the **node\_external\_reference.h** file. * To test fast APIs, run the tests in a loop with a decent iterations count to trigger V8 for optimization and to prefer the fast API over the slow one. * The fast callback must be idempotent up to the point where error and fallback conditions are checked because otherwise, executing the slow callback might produce visible side effects twice. ### Fallback to the slow path Fast APIs support fallback to the slow path (implementation that uses V8 internals) in case logically it is wise to do so, for example, when providing a more detailed error. The fallback mechanism can be enabled and changed from the caller JavaScript function and the fast API function declaration. Passing a `true` value to `fallback` option will force V8 to run the slow path with the same arguments. In V8, the options fallback struct is defined as **FastApiCallbackOptions** under the **v8-fast-api-calls.h** file. #### Example of C++ fallback ```cpp // Anywhere in the execution flow, you can set fallback and stop the execution. static double divide(const int32_t a, const int32_t b, v8::FastApiCallbackOptions& options) { if (b == 0) { options.fallback = true; return 0; } else { return a / b; } } ``` Anywhere in the execution flow, you can set fallback and stop the execution. ### Example usage * On JavaScript side: ```js title="Example usage of a javascript user-land function" const { divide } = internalBinding('custom_namespace'); ``` * On the C++ side: ```cpp title="Example module declaration in C++ with V8" #include "v8-fast-api-calls.h" namespace node { namespace custom_namespace { static void divide(const FunctionCallbackInfo& args) { Environment* env = Environment::GetCurrent(args); CHECK_GE(args.Length(), 2); CHECK(args[0]->IsInt32()); CHECK(args[1]->IsInt32()); auto a = args[0].As(); auto b = args[1].As(); if (b->Value() == 0) { return node::THROW_ERR_INVALID_STATE(env, "Error"); } double result = a->Value() / b->Value(); args.GetReturnValue().Set(result); } static double FastDivide(const int32_t a, const int32_t b, v8::FastApiCallbackOptions& options) { if (b == 0) { options.fallback = true; return 0; } else { return a / b; } } CFunction fast_divide_(CFunction::Make(FastDivide)); static void Initialize(Local target, Local unused, Local context, void* priv) { SetFastMethod(context, target, "divide", Divide, &fast_divide_); } void RegisterExternalReferences(ExternalReferenceRegistry* registry) { registry->Register(Divide); registry->Register(FastDivide); registry->Register(fast_divide_.GetTypeInfo()); } } // namespace custom_namespace } // namespace node NODE_BINDING_CONTEXT_AWARE_INTERNAL(custom_namespace, node::custom_namespace::Initialize); NODE_BINDING_EXTERNAL_REFERENCE( custom_namespace, node::custom_namespace::RegisterExternalReferences); ``` * Update **node\_external\_reference.h** Since our implementation used **int(const v8::FastApiCallbackOptions& options)** signature, we need to add it to external references if it is not available and in `ALLOWED_EXTERNAL_REFERENCE_TYPES`. ```cpp title="Example external reference declaration" using CFunctionCallbackReturningDouble = double (*)(const int32_t a, const int32_t b, v8::FastApiCallbackOptions& options); ``` --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/using-v8-fast-api-in-node-js-core Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Implementing Node.js URL parser in WebAssembly with Rust description: "A Fordham capstone: reimplementing Node's WHATWG URL parser in Rust and WebAssembly to cut C++ bridge cost." date: 2022-02-28 tag: performance series: url-parsing author: Yagiz Nizipli canonical: "https://www.yagiz.co/implementing-node-js-url-parser-in-webassembly-with-rust" markdown: "https://www.yagiz.co/implementing-node-js-url-parser-in-webassembly-with-rust.md" --- # Implementing Node.js URL parser in WebAssembly with Rust > A Fordham capstone: reimplementing Node's WHATWG URL parser in Rust and WebAssembly to cut C++ bridge cost. *Published: 2022-02-28 · Tag: performance* --- Even though, this started as an experiment, implementing the URL parser in Rust using WebAssembly became the graduation project for my Masters in Computer Science at Fordham University. ## A brief backstory I've started my Master's program on September 2021 and moved to New York from Istanbul, Turkey after working in the field for 10+ years. I've met with amazing people and professors in Fordham and took my graduation project at the end of my second semester (which is called Capstone project). The goal of the Capstone Project is to give students and future engineers to work on a project with the pricinciples of Software Engineering to prepare them for the field. Since, I've had the field experience, I convinced my advisor, William Lord, to select a project which will make a significant contribution to one of my favorite runtime environments, Node.js. ## Personal goal for graduation My main goal for selecting my graduation project at Fordham were; 1. The project should be technically challenging 2. I need to learn and experience a new technology 3. I need to give back to community. (If I'm going to spend more than 4 months on a project, I didn't want it to go to waste) 4. I want work on performance related stuff (which I didn't had any chance in the past, due to my small startup experience) ## The elevator pitch I saw an opportunity on the 48th page on Github issues on Node repository. There was a request for reimplementation for WHATWG URL Parser in WASM. **This was an important issue because it was mentioning:** * performance problems with C++ bridge and URL parser * "choose any technology you want" and compile into webassembly if it justifies the performance impact * one of the most used functions in Node.js with a potential of changing a lot of things * being 100% API compliant with the existing implementation So, here is my elevator pitch I've did to a bunch of students and teachers at Fordham University. ![WHATWG URL Parser](https://www.yagiz.co/content/whatwg-url-parser-slide-1.png) ![Wanted: Reimplement WHATWG URL Parser in WASM](https://www.yagiz.co/content/whatwg-url-parser-wanted-issue.png) ![WHATWG URL Parser Problem: C++ Bridge](https://www.yagiz.co/content/whatwg-url-parser-problem.png) ![WHATWG URL Parser Solution](https://www.yagiz.co/content/whatwg-url-parser-solution.png) ![Similar Studies](https://www.yagiz.co/content/whatwg-url-parser-similar-studies.png) ## Beginnings I've started reading the URL parser specification and understand the state machine behind it. Even though, I've implemented lots of state machines in the past, due to my selection of Rust (due to my eagerness to learn it), it was quite new for me to properly implement it using Rust. I've created a repository and implemented an initial PoC with scheme parser support on the URL side and the complete URLSearchParams implementation. ## Benchmarks Before starting to implement the URL Parser state machine in Rust, I wanted to see how I am doing and what is the impact of the work I was going to create. I started using `benchmarkify` and added simple benchmark to test my code. ```javascript title="Benchmark" const {URL: RustURL, URLSearchParams: RustURLSearchParams} = require('url-wasm') const Benchmarkify = require("benchmarkify"); const benchmark = new Benchmarkify("URL vs. Rust URL", {minSamples: 1000}).printHeader(); const index = benchmark.createSuite('URL') index.add('URL', () => new URL('https://www.google.com/path/to/something')) index.add('Rust::URL', () => new RustURL('https://www.google.com/path/to/something')) const search_params_set = benchmark.createSuite('URLSearchParams.set') search_params_set.add('URLSearchParams.set', () => { let searchParams = new URLSearchParams('hello=world') for (let i = 0; i < 100; i++) { searchParams.set(`key-${i}`, `value-${i}`) } return searchParams.toString() }) ``` The goal of the benchmark was to create a baseline before I did the serious stuff and see how it impacted through the releases I've did, and the progress I've made. ![Benchmark results](https://www.yagiz.co/content/whatwg-url-benchmark.png) Shockingly, the benchmarks were really bad. My Rust **URL implementation**, which didn't really do anything expect iterating through the input 1 time, was 27% slower than the actual implementation. Which was really shocking for me is that my **URLSearchParams** implementation was 86% adnd 95% slower for `set` and `append` functions, which basically just manipulated a vector inside Rust. The results from my benchmarks were really bad, and it was alarming. I was doing something wrong. ## Assumptions, assumptions, assumptions... Here are my assumptions before diving into it. 1. I was pretty sure I wasn't using Rust in a performant way. 2. My benchmarks were wrong, and v8 with JIT compiler was caching JavaScript in a much more performant way than WebAssembly 3. I was questioning myself and thought I've missed a really important bullet point in the implementation. 4. The javascript project auto-generated by `wasm-pack` was not performant. (This assumption was based on my experience with `auto-generated` libraries and how a general approach is always slower than a implementation-specific approach) ## Researching **1. I was pretty sure I wasn't using Rust in a performant way.** Upon my research I've realized that there are 2 different flags where I can put into `cargo.toml` to compile the Rust code in a performant way. One of them was `opt-level=3` which basically told the compiler to compile the code so that speed is preferred over space, and do it in the most agressive way. The second one was using `lto = true`. (LTO means Link-Time Optimization. It is generally set up to use the regular optimization passes used to produce object files) After adding the 2 options, **I saw 3-5% performance improvement** compared to previous one. It was not enough. **2. My benchmarks were wrong, and v8 with JIT compiler was caching JavaScript in a much more performant way than WebAssembly** I've did some research and saw [this amazing article](https://flaviocopes.com/node-runtime-v8-options/). I've did some experimenting with the flags but did not see an major improvement compared to the previous benchmark that would reason with the performance regression. This made me question if WebAssembly was the actual reason behind this issue. Upon researching and reading I've realized that UTF-16 to UTF-8 conversion between WebAssembly and Node.js using `TextEncoder` and `TextDecoder` was the primarily reason and the bottleneck. **3. I was questioning myself and thought I've missed a really important bullet point in the implementation.** The best way to know if you're on the correct path was to actually ask for people's help who have more experience on it than you. I've reached to [Matteo Collina](http://github.com/mcollina) and [James M. Snell](http://github.com/jasnell) from the Node.js Foundation Technical Steering Commitee. Upon talking with James, he suggested I should try to reimplement the JavaScript layer without the help from `wasm-pack` and see what would be the change in terms of performance. **4. The javascript project auto-generated by `wasm-pack` was not performant.** I've started reimplementing the JavaScript layer and just benchmarked the initialization (passing the string from JavaScript to WebAssembly (Rust)). ```javascript title="Here's the only code I've benchmarked on it:" class URL { #ptr = null constructor(url) { if (url) { const [ptr, vector_len] = passStringToWasm(url, instance.exports.__wbindgen_malloc, instance.exports.__wbindgen_realloc); this.#ptr = instance.exports.url_new(ptr, vector_len) } } } ``` ### Benchmark results ![Benchmark results](https://www.yagiz.co/content/whatwg-url-benchmark-results.png) Upon implementing it and asking guidance from Matteo (due to his experience in working with Buffers), I've made the implementation faster. (18% compared to 8% performance degregation). ## Conclusion I've realized that even though WebAssembly is really performant for certain CPU/GPU intensive tasks, it was not the correct technology for small operations including high string encoding and decoding. Companies such as 1Password prefer using WebAssembly over any other solution, is mainly because they're using CPU/GPU intensive tasks focusing on cryptography and encryption where JavaScript and v8 does not have the capacity. Even though this was an educational experience for me, I'll implement the WHATWG URL parser in JavaScript and continue my effort to make it more performant than the existing solution using v8. The Rust based implementation lives inside [https://github.com/anonrig/url](https://github.com/anonrig/url). --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/implementing-node-js-url-parser-in-webassembly-with-rust Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: "PostgreSQL: Optimizations and indexes" description: "PostgreSQL 14 index types, when they help, and how to read the query planner." date: 2022-02-26 tag: database author: Yagiz Nizipli canonical: "https://www.yagiz.co/postgresql-index-performance" markdown: "https://www.yagiz.co/postgresql-index-performance.md" --- # PostgreSQL: Optimizations and indexes > PostgreSQL 14 index types, when they help, and how to read the query planner. *Published: 2022-02-26 · Tag: database* --- Today, I'm going to talk about one of the most interesting topics in software and computer engineering, performance. PostgreSQL, as one of the most advanced database management systems out there, lacks the ability to suggest indexes according to your query planner. ## Indexes There's currently 6 different index types supported in [PostgreSQL 14](https://www.postgresql.org/docs/current/indexes-types.html). I'll briefly talk about these types and create a baseline for you to understand the advantages and disadvantages of using them. ```sql title="An example SQL statement to create a hash index" CREATE INDEX name ON table USING HASH (column); ``` ### B-tree Index * B-trees focuses on equality and range queries on data that can be sorted in some ordering. If you're using the following equality operators, PostgreSQL query planner will consider using a B-tree index. (`<`, `<=`, `=`, `>=`, `>`) * `BETWEEN`, `IN`, `IS NULL` or `IS NOT NULL` sql queries can also be implemented with a B-tree index search. If you're using `SELECT * FROM repositories WHERE created_at BETWEEN $1 AND $2` kind of queries, B-tree indexes are the correct choice for you. ### Hash Index * Hash indexes store a 32-bit hash code derived from the value of the column. Therefore, it's only used with `=` operator. ### GiST Index * Lossy Generalized Search Tree index. > A GiST index is __lossy_, meaning that the index might produce false matches, and it is necessary to check the actual table row to eliminate such false matches. (PostgreSQL does this automatically when needed.) GiST indexes are lossy because each document is represented in the index by a fixed-length signature. The signature is generated by hashing each word into a single bit in an n-bit string, with all these bits OR-ed together to produce an n-bit document signature. When two words hash to the same bit position there will be a false match. If all words in the query have matches (real or false) then the table row must be retrieved to see if the match is correct. `` CREATE INDEX __`name`_ ON __`table`_ USING GIST (__`column`_); `` ### SP-GiST Index * SP-GST is an abvreviation for space-partitioned GiST. SP-GiST supports partitioned search trees such as quad-trees, k-d trees and suffix trees. The common feature of these structures is that they repeatedly divide the earch space into partitions that need not be of equal size. * More appropriate explanation of GiST and SP-GiST can be found on [here](https://gis.stackexchange.com/questions/374091/when-to-use-gist-and-when-to-use-sp-gist-index). ### GIN Index * Generalized inverted indexes prefered mostly for text search. GIN indexes stores only the words as `tsvector` values. * A table check is required when using a query with weights. `` CREATE INDEX __`name`_ ON __`table`_ USING GIN (__`column`_); `` ### BRIN Index * Introduced on PostgreSQL 9.5 * BRIN index is the block-range index that allows serious performance boosts which involve BETWEEN queries. * 20 times better compared to B-Tree index. [Source.](https://www.percona.com/blog/2019/07/16/brin-index-for-postgresql-dont-forget-the-benefits/) `CREATE INDEX testtab_date_brin_idx ON testtab USING BRIN (date);` ## Index Recommendations While looking for a good index recommendation extension for PostgreSQL, I saw an extension developed by Powa team on [Github](https://github.com/powa-team/pg%5Fqualstats). ### Installation 1. Clone the repository on the preferred location using the following command: `git clone git@github.com:powa-team/pg_qualstats.git`. 2. Move to the correct folder using `cd pg_qualstats` and before running installation make sure that you have `postgres` in your environment. 3. After that, you need to run `make install` on the root directory. This will install the extension using the PostgreSQL header files and link it to your build. 4. Navigate to your `postgresql.conf` file and manipulate the line containing preload libraries to `shared_preload_libraries = 'pg_qualstats'` 5. Restart the PostgreSQL process ### Usage On the database where you want to get recommendations create the extension using the following SQL statement: `CREATE EXTENSION pg_qualstats;`. Now you're good to go! Execute different queries, preferably from your application and later query the following SQL statement: ```sql title="Example recommendation from pg_qualstats" SELECT v FROM json_array_elements( pg_qualstats_index_advisor(min_filter => 50)->'indexes') v ORDER BY v::text COLLATE "C"; v --------------------------------------------------------------- "CREATE INDEX ON public.adv USING btree (id1)" "CREATE INDEX ON public.adv USING btree (val, id1, id2, id3)" "CREATE INDEX ON public.pgqs USING btree (id)" (3 rows) ``` This query will give you possible indexes which can improve certain queries you've executed. For more detailed explanation on this awesome extension, please visit the extensions official [Github repository](https://github.com/powa-team/pg%5Fqualstats). --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/postgresql-index-performance Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Tracing query performance with Knex.js description: A small Knex.js helper that uses perf_hooks to time queries and print planner output so you can see why a query is slow. date: 2022-01-23 tag: performance author: Yagiz Nizipli canonical: "https://www.yagiz.co/tracing-query-performance-with-knex-js" markdown: "https://www.yagiz.co/tracing-query-performance-with-knex-js.md" --- # Tracing query performance with Knex.js > A small Knex.js helper that uses perf_hooks to time queries and print planner output so you can see why a query is slow. *Published: 2022-01-23 · Tag: performance* --- Over the years, I found myself searching for the same exact problem whenever I was using Knex.js, the query planner for Node.js. Since, the following code snippet is small enough to not be a Node library, I'm intrigued to share this as a blog post. Initially, a blog post from [Atomic Object](https://spin.atomicobject.com/2017/03/27/timing-queries-knexjs-nodejs/) proposed to use `query` and `query-response` events on the knex.js instance, since it's also an event emitter, but I guess it's time to use **perf\_hooks** and properly mark the performance of a query using the native tooling. ```javascript import { PerformanceObserver, performance } from 'perf_hooks' const performanceObserver = new PerformanceObserver((list) => { list.getEntries().forEach((entry) => { console.group(`${entry.duration.toFixed(2)} ms`) console.info(entry.name) console.groupEnd() }) }) performanceObserver.observe({ buffered: true, entryTypes: ['measure'] }) pg.on('query', (query) => { const id = query.__knexQueryUid performance.mark(`${id}-started`) }).on('query-response', (response, query) => { const id = query.__knexQueryUid performance.mark(`${id}-ended`) performance.measure(query.sql, `${id}-started`, `${id}-ended`) }) ``` Small snippet using perf_hooks ![Tracing query screenshot](https://www.yagiz.co/content/tracing-query-screenshot.png) With the following code, we'll see the following code on our terminal. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/tracing-query-performance-with-knex-js Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: "How to generate a valid and fast Sudoku board from scratch?" description: "How to generate a valid Sudoku board from scratch, and what “valid” actually requires." date: 2021-10-15 tag: algorithms author: Yagiz Nizipli canonical: "https://www.yagiz.co/sudoku-generating-valid-one" markdown: "https://www.yagiz.co/sudoku-generating-valid-one.md" --- # How to generate a valid and fast Sudoku board from scratch? > How to generate a valid Sudoku board from scratch, and what “valid” actually requires. *Published: 2021-10-15 · Tag: algorithms* --- I don't believe I'm writing a blog post related to Sudoku since it's already a solved problem/exercise and it's widely known, but I wanted to share my experience and my thought thinking about creating a valid Sudoku board and what it means to have a valid Sudoku. ## The Story Before diving into my obvious task in hand, I wanted to share the **before**. I'm currently having my Masters education in Computer Science program in Fordham University, New York. One of my classes, Software Engineering, gave us a homework about developing a Sudoku application for a imaginary company and what are the activity diagrams for a particular task: "Generating a Sudoku board". Before thinking about any algorithms, I wanted to give some information about what it means to generate a Sudoku board, and a valid one. ## Requirements 1. Board can be in different sizes. (Ex: 4x4, 9x9, 16x16) 2. Board should have a unique solution. If a Sudoku puzzle is solved with more than 1 approach, than it's not a Sudoku board, it's a puzzle? (lol) 3. Board should have some sort of understanding of what a difficulty is, or should have a basis of generating within a difficulty. (For the sake of simplicity, I won't be covering difficulty in this post) 4. Each board cell should be unique in their corresponding row, column and their square. ![With and without Sudoku solutions](https://www.yagiz.co/content/sudoku-with-solutions.png) An example Sudoku puzzle with a solution ## The Goal Our goal in this blog post is to create a thought process and a basic algorithm for generating a solvable Sudoku puzzle with a unique solution and have the flexibility to introduce difficulty level later on to the algorithm. I believe that the goal can be achieved in a reduced 2 step process. 1. Generating a full/solved Sudoku puzzle where every cell has a valid value and respects the board size. 2. Generating a uniquely solvable game board from a puzzle. ## Step 1: Generating a solution ### Glossary * **Square** is referenced to each square inside the Sudoku board. For a 4x4 Sudoku, there is 4 squares. For 9x9, there is 9, and etc. * **Board size** is referenced to the width of the Sudoku board. For 9x9, it is 9. ### Assumptions * The **board size** is an integer and have a value of more than 0. * The **square count** is an integer and have a value of more than 0. * **Board size divided by square count should equal to the square count**. (To preserve the validity of the Sudoku board) ### Thought process Let's generate a solution by iterating through each square in the Sudoku board (where for 4x4 there is 4 square, for 9x9 there is 9) for each number and fill it one by one while checking for the validity of the cell we're putting it. The following algorithm is an example of Breadth first search algorithm where each node has a validity and inserted into the tree according to the rule. For more information about the algorithm please look into the [Wikipedia article](https://en.wikipedia.org/wiki/Breadth-first%5Fsearch). ### Algorithm 1. Iterate x from 1 to board\_size 2. For each x, iterate from 0 to square\_count, y 3. Find all available positions in the given y. 4. Filter available positions according to the forbidden positions list. _(Forbidden position list is referred to an array of indexes where the current iteration tried but could not put a value due to a child having 0 valid positions for insertion)_ 5. Filter available positions according to my previous positions within the same x. _(This preserves the validity of the Sudoku within the same x)_ 6. Select a random position from the filtered list. If a position exist, update the board and continue to step 2\. If no positions are available for insertion, add the last value in previous positions list within the same x to the forbidden positions list, and re-iterate y-1 with the updated forbidden list. ![Sudoku solution steps](https://www.yagiz.co/content/sudoku-solution-steps.png) The visual representation of inserting numbers on 2x2 board ## Step 2: Generating a uniquely solvable solution ### Assumptions * An existing solution is available to use it as a base reproduced using Step 1. * Assume a set of difficulty ranges are available according to the minimum empty cells in a game. ### Algorithm 1. Select a random cell from the board. If the cell is not-empty and not-visited continue to step 2\. If the cell is visited move to step 6. 2. Remove the selected cell from the game board. 3. Calculate solutions for the removed cell on the board. If the uniqueness is achieved in solutions, continue to step 4\. If there are more than 1 solution to the given removed cell, undo the removal. 4. Set the current cell as visited. 5. Return to step 2. 6. Repeat step 2 until random and non-empty cell count is 0. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/sudoku-generating-valid-one Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Performance metrics and benchmarking on Node.js description: How to time Node.js functions with perf_hooks without fooling yourself. date: 2021-05-16 tag: performance author: Yagiz Nizipli canonical: "https://www.yagiz.co/performance-metrics-and-benchmarking-nodejs" markdown: "https://www.yagiz.co/performance-metrics-and-benchmarking-nodejs.md" --- # Performance metrics and benchmarking on Node.js > How to time Node.js functions with perf_hooks without fooling yourself. *Published: 2021-05-16 · Tag: performance* --- Lately, I've found myself worrying more and more about performance and the amount of time it takes for a function or a task to be taken on Node.js. Particularly, I wanted to benchmark and compare a particular code change to the original one. While researching for ways to implement and compare these functions, I came across Node.js's **perf_hooks** package which was added in the 8.5.0 release. For API definitions and use cases please refer to the [documentation](https://nodejs.org/api/perf%5Fhooks.html). ## Performance Measurement Let's first dive into a basic example of calculating the amount of time in milliseconds it takes to execute a specific function: ```js import { performance } from 'perf_hooks' async function hardToSwallowPills() { performance.mark('start-benchmarking') await doSomeAsyncTask() performance.mark('end-benchmarking') performance.measure('hardToSwallowPills', 'start-benchmarking', 'end-benchmarking') } ``` The important part you need to address in this code is, you need to manually mark a specific point in your codebase before and after the execution of it. After that, you need to measure the difference using **performance.measure** which takes the name/identifier of the measurement as the first parameter. In order to log these measurements, you'll need the PerformanceObserver implementation from the **perf_hooks**. ```js import { PerformanceObserver } from 'perf_hooks' const performanceObserver = new PerformanceObserver((list) => { list.getEntries().forEach((entry) => { logger .withTag('performance') .info(`${entry.name} took ${entry.duration.toFixed(2)} ms`) }) }) performanceObserver.observe({ entryTypes: ['measure'], buffered: true }) ``` ## Benchmarking In order to benchmark N number of implementations, you need to give the same input to both of them and compare them using the amount of time it takes to finish that task. For Node.js there's a really cool library called \`benchmark\` and provides a comparison and shows us the fastest implementation of our input. Example implementation is: ```js import Benchmark from 'benchmark' import { sort, randomArray } from './bubble.js' import { sort: quickSort } from './quick.js' const suite = new Benchmark.Suite() const random = randomArray(100, 5000000) suite .add('BubbleSort', () => sort(random)) .add('QuickSort', () => quick(random)) .on('cycle', function(event) { console.log(String(event.target)) }) .on('complete', function() { console.log('Fastest is ' + this.filter('fastest').map('name')) }) .run({ 'async': true }) ``` This will give you a brief result of your comparison like: ```bash ➜ algorithms node sorting/bubble.benchmark.js BubbleSort x 5,721,443 ops/sec ±0.23% (96 runs sampled) ``` --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/performance-metrics-and-benchmarking-nodejs Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: Timing Attacks on Node.js description: "How timing attacks show up in Node.js string comparison, and what eslint-plugin-security actually flags." date: 2021-03-24 tag: security author: Yagiz Nizipli canonical: "https://www.yagiz.co/timing-attacks-on-node-js" markdown: "https://www.yagiz.co/timing-attacks-on-node-js.md" --- # Timing Attacks on Node.js > How timing attacks show up in Node.js string comparison, and what eslint-plugin-security actually flags. *Published: 2021-03-24 · Tag: security* --- I've been working with Node.js for quite a long time. So, believe me when I say there's a library called **eslint-plugin-security** to detect common mistakes and security flaws you make while you're coding. Before discussing timing attacks on node.js, let's start by talking about my Eslint configuration on Socketkit's new privacy & security-oriented tracking API. With the release of Node 15, I started converting the API to modules and added/improved my usual stack by adding more and more Eslint rulesets to make my code more persistent for outsiders. (if that day comes) Here's my \`.eslintrc\` configuration for Socketkit's new Tracking API ```json title="Example .eslintrc configuration" { "extends": [ "plugin:prettier/recommended", "plugin:import/errors", "plugin:import/warnings", "plugin:security/recommended" ], "plugins": ["prettier", "import", "security"], "parserOptions": { "sourceType": "module", "ecmaFeatures": { "modules": true }, "ecmaVersion": 2020 }, "env": { "node": true, "es6": true }, "rules": { "import/extensions": ["error", "always", { "ignorePackages": true }] } } ``` I was in the middle of writing an endpoint for update password action for users of our web page. I didn't want to take the usual approach and use Ory Kratos' redirection loop based authentication flow and convert our browser based forget password flow to node.js based API flow recommended by Ory Kratos. Since our API was based on fastify, we handled our day to day validations using the presupported, awesome library called \`ajv\`. Here's an example schema by ajv and fastify, which checks for the type of 2 parameters: \`password\` and \`password\_again\` and the existence of it using the required parameter from ajv. ```javascript title="Example schema using AJV and Fastify" { schema: { body: { type: 'object', properties: { password: { type: 'string' }, password_again: { type: 'string' }, }, required: ['password', 'password_again'], }, response: { 200: { type: 'object', properties: { state: { type: 'boolean' }, }, }, }, }, } ``` Since, **password** and **password\_again** should be equal in order to make sure the user didn't want to change their password to a faulty one, we had to make sure that both of them is equal to each other. Our handler should have been this: ```javascript title="Example implementation" { handler: async ({ accounts: [account] }) => { if (password !== password_again) { throw new f.httpErrors.preconditionFailed(`Passwords should match.`) } return { state: true } }, } ``` Unfortunately, the following code produces an attack vector called \`Timing attack\`. When searched upon it's clear that \`eslint\` only checks for the name of the variable \`password\` and this is a false positive in terms of security. But for the sake of this article, let's continue on investigating the cause of it and try to improve our code. (Make it hacker proof) ![Hackerman](https://www.yagiz.co/content/hackerman.png) Hackerman ### Let's start by feeling like a hacker from now on In cryptography, a timing attack is a side-channel attack in which the attacker attempts to compromise a cryptosystem by analyzing the time taken to executive cryptographic functions. If we were calculating a SHA hash according to the password we got from the input, the execution time of calculating that particular hash would have been directly correlated with the length of the input. Additionally, information can leak from a system through measurement of the time it takes to respond to certain queries. (According to Wikipedia) In order to resolve this side-channel attack method, there is a specific function in \`crypto\` library in Node.js ### Here comes the **crypto.timingSafeEqual(a, b)** According to the fantastic Node.js contributors and developers, here's the definition of this function: > This function is based on a constant-time algorithm. Returns true if a is equal to b, without leaking timing information that would allow an attacker to guess one of the values. This is suitable for comparing HMAC digests or secret values like authentication cookies or capability URLs. So our following code should have been the following: ```javascript title="The correct implementation" { handler: async ({ accounts: [account] }) => { if ( !crypto.timingSafeEqual( Buffer.from(password), Buffer.from(password_again), ) ) { throw new f.httpErrors.preconditionFailed(`Passwords should match.`) } return { state: true } }, } ``` **P.S.:** Since we're not comparing two hash values or making any cryptographic calculations, this is absolutely unnecessary and a costly operation. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/timing-attacks-on-node-js Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing. --- title: "Cordova, React Native, Swift: What is really next?" description: "A 2020 comparison of Cordova, React Native, and native Swift for mobile apps." date: 2020-11-22 tag: coding author: Yagiz Nizipli canonical: "https://www.yagiz.co/cordova-react-native-swift" markdown: "https://www.yagiz.co/cordova-react-native-swift.md" --- # Cordova, React Native, Swift: What is really next? > A 2020 comparison of Cordova, React Native, and native Swift for mobile apps. *Published: 2020-11-22 · Tag: coding* --- I'm quite excited about writing about mobile application development. It has been in my mind for quite some time now. Before going into discussing my experience on these technologies I want to give a brief about my background. I'm a Software Architect and a Full Stack Developer who was working with JavaScript for the past 6 years. Since the old days of JavaScript with ES5, the language has grown a lot. It's quite obvious. But it's important to remember our past to understand our future. So, let's dive into the technologies! ## Cordova I don't think it's fair for me to write about Cordova since I'm not an active developer for the past 4 years. Back in the days, a simple animation using CSS3 Keyframe animation was not working as intended in Android devices. Since they were using the default browser and not the Chromium it was intended to run in. Since those days, Cordova has evolved and started to ship itself with a smaller version of chromium inside itself. Being able to run your website as a mobile application using Cordova was and still is a really interesting and fascinating feature for most of the web developers who just wanted to try the mobile application development. Additionally, for a team who focuses not only on a specific technology but on JavaScript and it's ecosystems, it's really an interesting investment to write your own PoC using Ionic (Angular). But the very thing it makes Cordova strong does also make the technology weak compared to others. The web. ### Main Issues * Browser layer of the application and the communication between the native code through it. * WKWebView being the only browser support for iOS devices and the inconsistencies between browsers (Android vs iOS). Due to constraints of the iOS platform, all browsers must be built on top of the WebKit rendering engine. * All communication must go through the HTML Rendering Engine (in iOS: WebView) and causes the performance of the application to have a strict correlation with the browser. ## React Native When React Native was first released it was quite a hype. The technology was unstable, but the promise it brings with it was quite fascinating! Before continuing to my thoughts on React Native, let's dive into the React Native ecosystem. > Up until October, Facebook was using its own fork of open-sourced React-Native. Since it came with extra pullbacks and effort, they switched to using the open sourced one. (You can see some issues where a Facebook developer doesn't encounter a specific bug, even though the same functionality was used inside Facebook itself, but because of Facebook using its own fork.) > Facebook releases a new version every month, up until September (I guess). After that, they started to release every 2 months (as far as I observe). Even the performance comparison with native apps showed a lot of promise both in React-Native and in JavaScript. Even in 2017, we wrote a production ready app all in React-Native. It was going great. Everybody was happy. The code was shipped fast, and since most of our developers had a JavaScript background, we were delivering on time and shipping production quality apps. Until the time I needed to write my own plugin to communicate with Swift side of the app through the JavaScript bridge, I was quite happy with it. Even, I suggested using it in my latest startup **SchoolApply**. The breaking changes in React Native caused my native code to misbehave due to unknown changes made by **react-native-git-upgrade** module which React Native and Facebook suggests to update the core itself. #### Main Issues * It's main dependency React and inconsistencies between the React-Native core and React. * React still doesn't follow the Semantic Versioning. From my perspective, software which is 4 years old, shouldn't be in 0.5X.XX. * Inconsistencies with native modules due to the breaking changes that happen to React-Native core a lot. (This is one of the main reason that React-Native defends of it being in 0.5X version, but this is again the main reason to use Semantic Versioning: to keep track of breaking changes and to inform developers about the ecosystem and deprecation) ## Swift It's quite unfair for me to judge a language by it's cover since I've been writing it for the past month but let's give it a shot. I had a quite prejudice since I thought React-Native or Cordova was doing the same exact thing as Swift: A mobile application. Every code written in Swift has a purpose. Yes, even the UITableView. It creates pain for new developers to get used to it and forces the user to use mainly XCode for development, even though there are solutions on the internet for using VSCode, Sublime or else. The auto layout constraints are quite hard for a new developer. There are really great libraries such as SnapKit, Cartography exists which aims to solve this ambiguity, but still not the default way of doing it. Even though all of the things above, the best experience I had as a developer is with Swift. There's no JavaScript bridge, WebView to get yourself bound into. You can use any library you want, with unlimited access to the device itself. --- Authored by Yagiz Nizipli Canonical: https://www.yagiz.co/cordova-react-native-swift Please attribute this content to Yagiz Nizipli and link back to the canonical URL when quoting or summarizing.