Predicates¶
Functions that inspect text and return boolean or structured results without modifying the input.
detect_scripts¶
detect_scripts ¶
detect_scripts(text: str) -> list[Script]
Return the set of Unicode scripts present in text, in order of first appearance.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> detect_scripts("Hello")
[Script.LATIN]
>>> detect_scripts("Hello Мир")
[Script.LATIN, Script.CYRILLIC]
inspect_auto_lang¶
inspect_auto_lang ¶
inspect_auto_lang(text: str) -> dict[str, str | list[str] | None]
Inspect how lang="auto" would resolve for the given text.
Use this to audit or log the detection decision made by the three-stage auto-detection pipeline.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> inspect_auto_lang("Київ")["chosen_lang"]
'uk'
>>> inspect_auto_lang("Москва")["reason"]
'script_default'
from disarm import inspect_auto_lang
inspect_auto_lang("Київ")
# {'script': 'Cyrillic', 'chosen_lang': 'uk', 'reason': 'discriminator', 'discriminators_hit': ['ї']}
inspect_auto_lang("Москва")
# {'script': 'Cyrillic', 'chosen_lang': 'ru', 'reason': 'script_default', 'discriminators_hit': []}
inspect_auto_lang("hello")
# {'script': None, 'chosen_lang': None, 'reason': 'no_detection', 'discriminators_hit': []}
See Language Detection for details.
is_mixed_script¶
is_mixed_script ¶
is_mixed_script(text: str) -> bool
True if text contains characters from more than one Unicode script.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> is_mixed_script("Hello")
False
>>> is_mixed_script("Hello Мир") # Latin + Cyrillic
True
has_bidi_conflict¶
has_bidi_conflict ¶
has_bidi_conflict(text: str) -> bool
True if text mixes strong left-to-right and strong right-to-left characters.
This is the precondition for Unicode Bidi display-reordering (UAX #9) — the
structural signal behind "BiDi Swap"-style spoofs, where an LTR brand label
sits beside an RTL domain (e.g. "varonis.com.ו.קום"). Unlike a
bidi-override (U+202x) check, it fires on the real letters: Latin /
Cyrillic / Greek / CJK are left-to-right; Hebrew / Arabic / Syriac / Thaana /
N'Ko are right-to-left; digits, punctuation and combining marks are neutral
and never create a conflict on their own.
A False result is not a safety guarantee.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> has_bidi_conflict("hello")
False
>>> has_bidi_conflict("helloא") # Latin + Hebrew
True
is_confusable¶
is_confusable ¶
is_confusable(text: str, *, target_script: str = 'latin', greedy: bool | None = None, preferred_aliases: list[str] | None = None) -> bool
True if text contains characters confusable with target-script characters.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> is_confusable("pаypal") # Cyrillic а looks like Latin a
True
>>> is_confusable("paypal") # all genuine Latin
False
unmapped_confusables¶
unmapped_confusables ¶
unmapped_confusables(*, target_script: str = 'latin') -> frozenset[str]
Every upstream confusable source disarm's bundled table does not fold (#563).
Read this as exposure, not as a score. A tool at 95% per-source coverage is not 95% safe — it is one query away from the other 5%, and this set is where an adaptive attacker goes when the mapped sources stop working.
Most of the set is out of scope rather than missing: a source whose upstream target
is non-Latin has no business in the to-Latin table. Cross-reference
:data:disarm.CONFUSABLES_VERSION and docs/provenance.md before reading any one
codepoint as a defect.
The set includes five ASCII characters — %, 0, 1, I and m. TR39
is a skeleton transform (m→rn, I/1→l, 0→O), so those are upstream sources; disarm
does not apply those rows because folding a legitimate ASCII m to rn corrupts
prose. Nothing is filtered out here: a coverage report that quietly drops rows reads
as coverage it does not have.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> unmapped = unmapped_confusables()
>>> "а" in unmapped # Cyrillic а IS folded, so it is not exposure
False
>>> "m" in unmapped # TR39 skeleton source m→rn, deliberately not applied
True
find_unmapped_confusables¶
find_unmapped_confusables ¶
find_unmapped_confusables(text: str, *, target_script: str = 'latin') -> list[tuple[str, int]]
Find confusable sources in text that disarm's table does not fold (#563).
The confusables analogue of :func:find_untranslatable, and it follows the same
convention: (character, byte_offset) pairs in order of appearance. This is what
turns :func:unmapped_confusables from a global number into something answerable
against your own traffic.
Composition runs exactly as it does in :func:normalize_confusables, so a
decomposed homoglyph whose precomposed form is mapped counts as covered rather
than as a gap — otherwise the report would disagree with what the transform does.
Offsets are anchored in text, never in the composed intermediate.
Ordinary English will report the letter m; see :func:unmapped_confusables for
why that is deliberate.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> find_unmapped_confusables("pаypal") # Cyrillic а folds — covered
[]
>>> find_unmapped_confusables("hello")
[]
is_ascii¶
is_ascii ¶
is_ascii(text: str) -> bool
True if all characters are in U+0000–U+007F.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> is_ascii("hello 123")
True
>>> is_ascii("café")
False
is_normalized¶
is_normalized ¶
is_normalized(text: str, *, form: NormalizationForm = 'NFC') -> bool
True if text is already in the specified normalization form.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> is_normalized("café") # NFC by default
True
>>> is_normalized("e\u0301", form="NFC") # NFD decomposed
False
is_zalgo¶
is_zalgo ¶
is_zalgo(text: str, *, threshold: int = 3) -> bool
Detect whether text contains zalgo-style combining mark abuse.
Returns True if any base character has more than threshold
consecutive combining marks in NFD decomposition.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> is_zalgo("café")
False
>>> is_zalgo("Việt Nam")
False
>>> is_zalgo("ḧ̸̡̢̧̛̗̱̜̼̯̞̙́̑̾̊̿̏̒̓̕ě̵̢̧̛̗̱̜̼̯̞̙̈́̑̾̊̿̏̒̓̕l̸̡̢̧̛̗̱̜̼̯̞̙̈́̑̾̊̿̏̒̓̕l̸̡̢̧̛̗̱̜̼̯̞̙̈́̑̾̊̿̏̒̓̕ơ̵̢̧̗̱̜̼̯̞̙̈́̑̾̊̿̏̒̓̕")
True
from disarm import is_zalgo
is_zalgo("café") # False (1 combining mark — normal)
is_zalgo("Việt Nam") # False (2 combining marks — normal)
# Zalgo: 'a' with 20 stacked combining graves
is_zalgo("a" + "\u0300" * 20) # True
is_suspicious_hostname¶
Renamed from is_safe_hostname in 0.9.1 — with the boolean inverted
If you are upgrading from is_safe_hostname, the return value's polarity was flipped
(safe → suspicious); a mechanical rename silently reverses your allow/deny branch.
See the Upgrading guide.
is_suspicious_hostname ¶
is_suspicious_hostname(hostname: str, *, contractions: bool = False) -> tuple[bool, HostnameAnalysis]
Flag a hostname as suspicious for Unicode homoglyph spoofing.
Returns (suspicious, analysis) where analysis is a
HostnameAnalysis with attributes:
suspicious: bool — True if a problem was detected (mixed-script, a bundled-table confusable, or a bidi-direction conflict). Because the confusable check is an any-character screen, this flags essentially every hostname with a non-Latin letter — legitimate (москва.рф) as well as spoofs — so it is a maximally conservative screen, not a precise verdict.scripts: list[str] — Unicode scripts found across all labels.mixed_script: bool — True if any single label contains more than one script.has_confusables: bool — True if confusable homoglyphs found.bidi_conflict: bool — True if the decoded hostname mixes strong left-to-right and strong right-to-left characters (the "BiDi Swap" reorder precondition). Folded intosuspicious.bidi_control: bool — True if the decoded hostname carries a UAX #9 bidi control character: an override (U+202D/U+202E), embedding (U+202A–U+202C), isolate (U+2066–U+2069) or directional mark (U+200E/U+200F/U+061C). Disjoint frombidi_conflict, which reads strong-direction letters only and is therefore blind to the RLO extension spoof. IDNA2008 disallows every character in the set, so this is folded intosuspiciousand the characters are stripped fromcanonical.has_invisible: bool — True if the decoded hostname carries a zero-width or invisible-format character:U+200B-U+200D,U+2060-U+2064,U+FEFForU+180E. Disjoint frombidi_control— these carry no direction at all, so neither bidi field can see them. Folded intosuspicious. They are removed before any other field is computed, so a hostname whose only non-ASCII is an invisible no longer reports a phantom script (U+FEFFsits in the Arabic Presentation Forms block).cross_label_script: bool — True if the labels span more than one distinct script. Broader and noisier thanbidi_conflict(it fires on benign IDN ccTLDs likegoogle.рф), so it is not folded intosuspicious; exposed for caller policy.label_scripts: list[list[str]] — per-label resolved scripts, left to right.whole_script_confusable: bool — True if any label is a whole-script confusable: single-script, non-Latin, whose confusable skeleton is entirely Latin (e.g. Cyrillicаррӏе→apple). A graded signal, not a verdict — on its own it fires on short non-Latin ccTLDs (ру→py) and on real words (оса→oca), so it is not folded intosuspicious.label_whole_script_confusable: list[bool] — per-label flags, parallel tolabel_scripts, so a caller can exclude the TLD label. The precise, low-false-positive policy iswsc(non-TLD label) and TLD-is-Latin(plus a caller-supplied protected-name list for the irreducibleоса-style case).canonical: str — Latin-normalized form of the hostname.
A hostname is flagged suspicious if any single label is mixed-script
(draws on more than one Unicode script, excluding Common/Inherited),
contains confusable homoglyphs, or has a bidi-direction conflict
(bidi_conflict), carries a bidi control character (bidi_control), or
carries a zero-width/invisible character (has_invisible).
The mixed-script rule is conservative and fails closed:
it flags benign combinations such as Latin+CJK as well as spoofing ones, so a
caller wanting a more permissive policy can inspect the mixed_script and
scripts fields and decide for itself.
A False (not-suspicious) result is not a safety guarantee. It means
only that no mixed-script label and no confusable from the bundled TR39
table was found. Confusables outside the bundled table are not detected and
report not-suspicious. Base allow/deny decisions on the granular findings
(including whole_script_confusable) plus your own policy — a detector can
attest the presence of a problem, never the absence of all problems.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> suspicious, analysis = is_suspicious_hostname("google.com")
>>> suspicious
False
>>> analysis.canonical
'google.com'
>>> _s, a = is_suspicious_hostname("arnazon.com", contractions=True)
>>> a.canonical
'amazon.com'
HostnameAnalysis¶
The second element of the tuple returned by is_suspicious_hostname():
| Attribute | Type | Description |
|---|---|---|
suspicious |
bool |
True if any label is mixed-script, contains a Latin-confusable character, or the hostname has a bidi-direction conflict, a bidi control character, or a zero-width/invisible character. An any-character confusable screen — it flags essentially every non-Latin hostname, so it is a maximally conservative screen, not a precise verdict |
scripts |
list[str] |
Unicode scripts found across all labels |
mixed_script |
bool |
True if any single label contains more than one script |
has_confusables |
bool |
True if any label contains a Latin-confusable character |
bidi_conflict |
bool |
True if the decoded hostname mixes strong LTR and RTL characters (the "BiDi Swap" precondition); folded into suspicious |
bidi_control |
bool |
True if the decoded hostname carries a UAX #9 bidi control character — override (U+202D/U+202E), embedding (U+202A–U+202C), isolate (U+2066–U+2069) or directional mark (U+200E/U+200F/U+061C). Disjoint from bidi_conflict, which reads strong-direction letters only. Folded into suspicious; the characters are stripped from canonical |
has_invisible |
bool |
True if the decoded hostname carries a zero-width or invisible-format character — U+200B–U+200D, U+2060–U+2064, U+FEFF, U+180E. Disjoint from bidi_control: these carry no direction at all. Folded into suspicious, and removed before any other field is computed, so they never reach scripts, mixed_script or canonical |
cross_label_script |
bool |
True if the labels span more than one script; broader/noisier than bidi_conflict (fires on benign IDN ccTLDs like google.рф), so not folded into suspicious |
label_scripts |
list[list[str]] |
Per-label resolved scripts, left to right |
whole_script_confusable |
bool |
True if any label is single-script, non-Latin, whose confusable skeleton is entirely Latin (аррӏе→apple). A graded signal, not a verdict — not folded into suspicious (fires on ру→py, оса→oca) |
label_whole_script_confusable |
list[bool] |
Per-label whole-script-confusable flags, parallel to label_scripts (exclude the TLD label for the precise policy) |
canonical |
str |
Latin-normalized form of the hostname |
from disarm import is_suspicious_hostname
suspicious, analysis = is_suspicious_hostname("google.com")
# suspicious = False, analysis.canonical = "google.com"
suspicious, analysis = is_suspicious_hostname("gооgle.com") # Cyrillic о's
# suspicious = True, analysis.mixed_script = True, analysis.has_confusables = True
# Whole-script spoof: an all-Cyrillic label whose skeleton is Latin
suspicious, analysis = is_suspicious_hostname("аррӏе.com")
# analysis.whole_script_confusable = True
# analysis.label_whole_script_confusable = [True, False] # spoof label, then the TLD
# analysis.canonical = "apple.com"
suspicious is a maximally conservative screen: because the confusable check is an any-character test and the most frequent Cyrillic/Greek letters are TR39 confusables, it flags essentially every non-Latin hostname — москва.рф as readily as аррӏе.com. A not-suspicious result is not a safety guarantee, and a suspicious one is not a precise verdict. For whole-script spoofs, use whole_script_confusable / label_whole_script_confusable: the precise, low-false-positive policy is whole_script_confusable(non-TLD label) ∧ (TLD is Latin/ASCII), applied by the caller — disarm deliberately does not model registrable boundaries (no PSL), and the irreducible оса-style case (a real word that skeletons to Latin) needs a caller-supplied protected-name list. See the Threat Model.