Support
Loading...

ICANN Releases Final String Similarity Evaluation Data and Guidelines for the 2026 New gTLD Round

31 July 2026

ICANN has published the finalized String Similarity Evaluation (SSE) Data together with the final String Similarity Evaluation Guidelines for the New gTLD Program: 2026 Round. These resources are intended to support the evaluation of applied-for top-level domain strings while helping minimize the possibility of user confusion caused by visually similar domain extensions. The documents were finalized following the completion of the public comment process.

The String Similarity Evaluation is a mandatory stage of the New gTLD Program: 2026 Round, as specified in Section 7.10 of the Applicant Guidebook. The evaluation framework is based on recommendations from the Generic Names Supporting Organization’s Final Report on the New gTLD Subsequent Procedures Policy Development Process and the Phase 1 Final Report of the Internationalized Domain Names Expedited Policy Development Process.

How the Evaluation Is Performed

During the evaluation, applied-for strings and their variant strings are examined to determine whether they are visually similar to existing or competing top-level domains. The assessment combines automated analysis with expert manual review. An SSE tool first generates a pre-screening report using the published similarity data, after which the String Similarity Evaluation panel completes a detailed review by applying the finalized guidelines.

The published guidelines describe how evaluators should conduct manual assessments and explain the role of the pre-screening report throughout the evaluation process. They also include an appendix outlining the complete operational workflow used by the SSE tool, helping ensure that all evaluations are performed in a consistent manner.

Finalized Data Covers the Entire RZ-LGR Repertoire

Alongside the guidelines, ICANN has also released the finalized SSE Data, which identifies pairs of code points considered visually similar by script experts and assigns an appropriate similarity level to each pair. The dataset covers the full repertoire of the Root Zone Label Generation Rules (RZ-LGR), enabling consistent evaluation of domain strings across a wide range of writing systems.

To support both technical implementation and human review, the SSE Data is available in machine-readable XML format for integration with the SSE tool, as well as in a human-readable HTML format for reference. ICANN has also published the complete data package, giving applicants and the broader Internet community access to the same technical resources used during string similarity evaluations.

Share this article:
Ask Jexi