ShinrAIHosted on STACKIT
Documentation

Languages and data categories

See what ShinrAI 1.5 recognises, how coverage is measured and which API surface to use.

15 model locales

One multilingual model, with regional Portuguese variants and Hebrew in beta. Website languages and API language codes are separate.

de

German

Available

95,5% strict-span F1

en

English

Available

97,3% strict-span F1

ja

Japanese

Available

98,5% strict-span F1

fr

French

Available

96,4% strict-span F1

es

Spanish

Available

97,6% strict-span F1

it

Italian

Available

96,1% strict-span F1

pl

Polish

Available

93,5% strict-span F1

pt-PT

Portuguese · Portugal

Available

94,6% strict-span F1

pt-BR

Portuguese · Brazil

Available

98,2% strict-span F1

ru

Russian

Available

96,9% strict-span F1

uk

Ukrainian

Available

95,0% strict-span F1

tr

Turkish

Available

93,2% strict-span F1

ar

Arabic

Available

90,1% strict-span F1

ko

Korean

Available

95,2% strict-span F1

he

Hebrew

Beta

71,2% strict-span F1

Business, medical and administrative documents; 200 texts per locale. Strict entity spans. Hebrew beta is included. Model-only evaluation by Innovius, not a service-wide accuracy guarantee.

View methodology and results →
Model locales, API codes and website languages

The website is available in English, German, Japanese, French and Spanish. This does not limit the model to those five languages.

The documented Azure text adapter accepts these language codes:

de en fr es it pl pt ru uk tr ar he ja ko

Portuguese uses one API language code but has two evaluated regional model locales. An accepted code does not guarantee every entity category or equal quality.

Read the adapter limits →

21 data categories

Examples are illustrative. Detection depends on the surrounding text and configuration.

People and organisations

People

Emma WeberPERSON

Organisations

Example GmbHCOMPANY

Contact and location

Cities

BerlinCITY

Street addresses

Musterstraße 12STREET_ADDRESS

Email addresses

emma@example.comEMAIL

Phone numbers

+49 30 000000PHONE

Postal codes

10115POSTAL_CODE

Finance

Bank accounts

DE89 3704 0044 0532 0130 00BANK_ACCOUNT

Crypto wallets

0x…CRYPTO_WALLET

Payment cards

4242 4242 4242 4242CREDIT_CARD

Personal identifiers

Dates of birth

14.03.1985DATE_OF_BIRTH

National identifiers

Tax / identity / social-security IDNATIONAL_ID

Vehicle plates

B AB 1234LICENSE_PLATE

Customer and case IDs

CUSTOMER-0042CUSTOMER_ID

Dates and ages

Dates

12 March 2026DATE

Ages

42 yearsAGE

Digital and confidential information

Usernames

example\emmaUSERNAME

Web addresses

https://example.comURL

Network addresses

192.0.2.1IP_ADDRESS

Secrets and access tokens

API_KEY_EXAMPLESECRET

Project codenames

Project AuroraPROJECT_CODENAME
Why 21 categories and 29 classes?

The model has 21 detection heads. People, cities, streets and organisations each have three subtypes; the other 17 categories have one. That gives 29 classes, not 29 unrelated data categories.

PERSON
common, uncommon, rare
CITY
major, medium, small
STREET_ADDRESS
generic, specific, local
COMPANY
international, national, regional

Origin, name form and other attributes help choose appropriate replacements. They are additional attributes, not additional languages or guaranteed detections.

Enterprise detectors and API-specific coverage

Enterprise adds deterministic rules and validation for supported identifiers. These are not included in the model-only scores above.

Adapter examples include US social-security numbers, bank routing numbers, SWIFT codes and vehicle identification numbers. Coverage and required context vary by API.

Compatibility adapters expose their documented categories and operations. Unsupported options are rejected; unmapped findings can produce warnings. Use the native API when you need findings outside an adapter mapping.

Native API and compatibility documentation →