Catalog
Korean AI datasets, by name
41–80 of 286 · page 2 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 40GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 380GB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 480MB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 190GB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 3.5GB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 180MB | AI Hub terms | Commercial use may be conditional |
| Ko-WinoGrande | Korean Instruction Tuning Dataset | Hugging Face (thunder-research-group) | 4.2MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Thunder-KoNUBench | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | 5.4MB | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 520GB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 450MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 310MB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 120hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 60hours | AI Hub terms | Commercial use may be conditional |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | Hugging Face (cfpark00) | 5.5MB | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| Smilegate UnSmile | Korean NLP Dataset | GitHub | Small | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 80hours | AI Hub terms | Commercial use may be conditional |
| TTS | Korean Speech Synthesis Dataset | AI Hub | 90hours | AI Hub terms | Commercial use may be conditional |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub (bab2min) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 220hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 150hours | AI Hub terms | Commercial use may be conditional |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 50hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 8GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 30GB | AI Hub terms | Commercial use may be conditional |
| Nemotron | Korean NLP Dataset | Hugging Face (nvidia) | 120MB | CC BY 4.0 | Commercial use may be conditional |
| FineWeb-Edu | Korean Pretraining Corpus | Hugging Face (eliceai) | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 1.6TB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 240MB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub | 180hours | AI Hub terms | Commercial use may be conditional |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 2.4GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 8TB | AI Hub terms | Commercial use may be conditional |
| OCR | Korean OCR Dataset | AI Hub | 290GB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | 18MB | License unknown | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 180MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 90MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 620MB | AI Hub terms | Commercial use may be conditional |