Catalog
Korean AI datasets, by name
121–160 of 286 · page 4 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| SMS | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| KorQuAD | Korean QA Dataset | KorQuAD | Large ( ) | CC BY-ND 4.0 | Commercial use allowed (catalog) |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | Large | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| IoT | Korean Speech Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | 120MB | License unknown | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Meme | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| KoSafetyBench | Korean LLM Benchmark Dataset | Hugging Face | 15MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean Toxicity Dataset | Korean Toxicity Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean Summarization Dataset | Korean Summarization Dataset | Hugging Face | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | 764MB | CC BY-SA 4.0 | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | GitHub | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | Hugging Face (Nagase-Kotono) | Large (7.85 GB) | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 45MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |