TECHNOLOGY
Technology & Data
Our core asset is Korean language education data, normalized into a form an AI model can learn from directly. See how we built it.
THE PROBLEM
We started out generating lesson plans with GPT-4o
The more lesson plans we generated, the further they drifted from the learners’ proficiency levels. So we measured it ourselves against the National Institute of Korean Language standard curriculum.
The reason was simple: no existing dataset included reliable proficiency-level classifications. So we built one.
SOURCE
Research papers, textbooks, the
national curriculum. That is where we began
Assembled domain by domain by researchers trained in Korean language education
STEP 1 · VOCABULARY
0entriesVocabulary
All levels 1 to 6 · 26,586 instances · 26,586 semantic classes
STEP 2 · GRAMMAR
0rowsGrammar
755 standard forms + 24,452 inflections · validated automatically by our patented method
STEP 3 · TOPIC
0categoriesTopics
2,775 unit mappings · 343 level-sequence entries
STEP 4 · FUNCTION
0categoriesFunctions
Communicative function categories · 5,610 unit mappings · 205 level-sequence entries
STEP 5 · CROSS MAPPING
0rowsCross-mapping
Vocabulary × grammar, vocabulary × function, vocabulary × topic, function × grammar and more. Separate lists become one graph
TOTAL
0rowsNormalized Korean education data
INTELLECTUAL PROPERTY
Intellectual property
Method for evaluating pronunciation accuracy based on Korean phoneme similarity
No. 10-2966195
Registered 2026.05.13
Method and system for automated validation of Korean grammar data
No. 10-2999681
Registered 2026.07.30
MCP SERVER
We deliver connected data to our products
Korean language data becomes useful to real learners only when all domains are connected. We serve that data through an MCP server and build products for non-native speakers of Korean on top of it.
Anything else you would like to know?
We welcome enquiries about partnerships and adoption.
Contact us
