Collection
Korean LLM training datasets
63 catalog items · GearDel
GearDel lists 63 catalog items in this group (63 text). Sources: 30 Hugging Face, 6 AI Hub, 10 GitHub. Examples: Korean Summarization Dataset, KULLM-v2, KoAlpaca-v1.1a. 46 are marked commercial-use-allowed in the catalog. Use this list to compare license and size, then open the original provider. GearDel does not host files.
Compare license and source before download. 46 of 63 carry a commercial-allowed catalog mark.
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | Large | AI Hub terms | Commercial use may be conditional |
| KULLM-v2 | Korean Instruction Tuning Dataset | Hugging Face | 180MB | Apache 2.0 | Commercial use allowed (catalog) |
| KoAlpaca-v1.1a | Korean Instruction Tuning Dataset | Hugging Face | Large | License unknown | Commercial use allowed (catalog) |
| DKTC (Dataset of Korean Threat Dialogue) | Korean Dialogue Dataset | GitHub | 1.8MB | CC BY-NC-SA 4.0 | Commercial use restricted |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | Large | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 3.5MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 24.5MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 1.2GB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 520MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 850MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 120MB | NIKL terms | Commercial use may be conditional |
| KoGPT2 | Korean Pretraining Corpus | Hugging Face | 3.2GB | License unknown | Commercial use allowed (catalog) |
| SNS | Korean Dialogue Dataset | AI Hub | 980MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 2.4GB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 3.5GB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 520MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 310MB | AI Hub terms | Commercial use may be conditional |
| TinyStories | Korean Translation Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| QARV | Korean Instruction Tuning Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Orca DPO | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | GitHub | Small | Apache 2.0 | Commercial use allowed (catalog) |
| LIMA | Korean Instruction Tuning Dataset | Hugging Face | Small | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face (devngho) | Large | MIT | Commercial use allowed (catalog) |
| FineWeb-Edu | Korean Pretraining Corpus | Hugging Face (eliceai) | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face (devngho) | Large | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face (coastral) | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Function Calling | Korean Instruction Tuning Dataset | Hugging Face (gyung) | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | GitHub | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| QG | Korean Instruction Tuning Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| WanJuanSiLu | Korean Dialogue Dataset | Hugging Face (opendatalab) | 124 GB | CC BY 4.0 | Commercial use allowed (catalog) |
| Role-Playing | Korean Instruction Tuning Dataset | Hugging Face (huggingface-KREW) | 28MB | Apache 2.0 | Commercial use allowed (catalog) |
| Ko-WinoGrande | Korean Instruction Tuning Dataset | Hugging Face (thunder-research-group) | 4.2MB | CC BY 4.0 | Commercial use allowed (catalog) |
| SmileStyle | Korean Pretraining Corpus | GitHub (smilegate-ai) | 35MB | MIT | Commercial use allowed (catalog) |
| KoCoSa | Korean Instruction Tuning Dataset | GitHub | 18MB | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| KoELECTRA | Korean Dialogue Dataset | GitHub (monologg) | Large (34GB) | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Ko-Llama | Korean Instruction Tuning Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| KIT-19 | Korean Instruction Tuning Dataset | Hugging Face | 85MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Korean Open-ORPO SFT+DPO | Korean Instruction Tuning Dataset | Hugging Face | 65MB | Apache 2.0 | Commercial use allowed (catalog) |
| KoAlpaca v1.0 Alpaca | Korean Instruction Tuning Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KOpen-Platypus | Korean Instruction Tuning Dataset | Hugging Face | 15.9MB | CC BY 4.0 | Commercial use allowed (catalog) |
| OpenAssistant Guanaco | Korean Instruction Tuning Dataset | Hugging Face | 16.7MB | Apache 2.0 | Commercial use allowed (catalog) |
| KcBERT | Korean Pretraining Corpus | Hugging Face | 7.90 GiB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| OpenOrca-KO | Korean Instruction Tuning Dataset | Hugging Face | 21.8MB | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | 764MB | CC BY-SA 4.0 | Commercial use may be conditional |
| xP3x | Korean Instruction Tuning Dataset | Hugging Face | 4,642,468rows | Apache 2.0 | Commercial use may be conditional |