Catalog
Korean AI datasets, by name
241–280 of 286 · page 7 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| KOpen | Korean Translation Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| KOpen-Platypus | Korean Instruction Tuning Dataset | Hugging Face | 15.9MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Korean Open-ORPO SFT+DPO | Korean Instruction Tuning Dataset | Hugging Face | 65MB | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Receipts OCR | Korean OCR Dataset | Hugging Face | 95MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Korean SAT Math | Korean NLP Dataset | Hugging Face (Quadyun) | 8.5MB | MIT | Commercial use allowed (catalog) |
| Korean Tourist Spot Dataset | Korean Computer Vision Dataset | GitHub | 8.4GB | License unknown | Commercial use allowed (catalog) |
| Korean-Light-OCR | Korean OCR Dataset | GitHub | 4.2GB | License unknown | Commercial use allowed (catalog) |
| KorMedMCQA QA | Korean QA Dataset | Hugging Face | 2.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KorNAT QA | Korean QA Dataset | Hugging Face | 2.6MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KorWikiTableQuestions | Korean QA Dataset | GitHub | 70K QA | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KoSBi | Korean NLP Dataset | GitHub | Large | MIT | Commercial use allowed (catalog) |
| Entailment | Korean NLP Dataset | GitHub | 6.2MB | Apache 2.0 | Commercial use allowed (catalog) |
| KoSimpleQA QA | Korean QA Dataset | Hugging Face | 301KB | License unknown | Commercial use may be conditional |
| KoSum | Korean Speech Dataset | Hugging Face (iontail) | 45GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| KOTE | Korean Sentiment Analysis Dataset | Hugging Face (searle-j) | Medium | MIT | Commercial use allowed (catalog) |
| KPoEM | Korean Sentiment Analysis Dataset | GitHub | 4.2MB | MIT | Commercial use allowed (catalog) |
| KRETA (Korean Rich Text Visual QA) | Korean OCR Dataset | Hugging Face | 125MB | License unknown | Commercial use allowed (catalog) |
| Role-Playing | Korean Instruction Tuning Dataset | Hugging Face (huggingface-KREW) | 28MB | Apache 2.0 | Commercial use allowed (catalog) |
| KsponSpeech | Korean Speech Recognition Dataset | AI Hub | 1000hours | AI Hub terms | Commercial use may be conditional |
| KSS Dataset | Korean Speech Synthesis Dataset | Hugging Face | 12hours+ | CC BY-NC-SA 4.0 | Commercial use restricted |
| KULLM-v2 | Korean Instruction Tuning Dataset | Hugging Face | 180MB | Apache 2.0 | Commercial use allowed (catalog) |
| LIMA | Korean Instruction Tuning Dataset | Hugging Face | Small | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| MeloTTS | Korean Speech Synthesis Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| OLKAVS | Korean Speech Recognition Dataset | GitHub | 1,150hours | License unknown | Commercial use may be conditional |
| ko-en | Korean Translation Dataset | GitHub | Large | License unknown | Commercial use allowed (catalog) |
| OpenAssistant Guanaco | Korean Instruction Tuning Dataset | Hugging Face | 16.7MB | Apache 2.0 | Commercial use allowed (catalog) |
| OpenOrca-KO | Korean Instruction Tuning Dataset | Hugging Face | 21.8MB | MIT | Commercial use allowed (catalog) |
| Orca DPO | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Pansori TEDxKR | Korean Speech Recognition Dataset | GitHub | ~3hours | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| ParaKQC | Korean NLP Dataset | GitHub | 10,000 (v1) | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| ko | Korean NLP Dataset | GitHub | 5.1MB | Other | Commercial use may be conditional |
| QARV | Korean Instruction Tuning Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| skt KVQA | Korean Computer Vision Dataset | Hugging Face | 100,445 | Korean VQA License | Commercial use restricted |
| SmileStyle | Korean Pretraining Corpus | GitHub (smilegate-ai) | 35MB | MIT | Commercial use allowed (catalog) |
| SNS | Korean Dialogue Dataset | AI Hub | 980MB | AI Hub terms | Commercial use may be conditional |
| SNU Ko-MuSR | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | Small | MIT | Commercial use allowed (catalog) |
| Speech Emotion Dataset (ko) | Korean Speech Dataset | GitHub | 3.2GB | License unknown | Commercial use allowed (catalog) |
| SQuARe | Korean NLP Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Table-VQA-ko | Korean OCR Dataset | Hugging Face | 185MB | Apache 2.0 | Commercial use allowed (catalog) |
| Tatoeba (ko-en) | Korean Translation Dataset | GitHub | 3.4MB | CC BY 2.0 | Commercial use allowed (catalog) |