Catalog
Korean AI datasets, by name
281–286 of 286 · page 8 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| TinyStories | Korean Translation Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Theory of Mind | Korean NLP Dataset | GitHub | 8.5MB | Apache 2.0 | Commercial use allowed (catalog) |
| ko | Korean QA Dataset | GitHub | 18.5MB | Apache 2.0 | Commercial use allowed (catalog) |
| WanJuanSiLu | Korean Dialogue Dataset | Hugging Face (opendatalab) | 124 GB | CC BY 4.0 | Commercial use allowed (catalog) |
| xP3x | Korean Instruction Tuning Dataset | Hugging Face | 4,642,468rows | Apache 2.0 | Commercial use may be conditional |
| Zeroth | Korean Speech Recognition Dataset | GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |