Collection
Korean speech datasets
38 catalog items · GearDel
GearDel lists 38 catalog items in this group (38 speech). Sources: 2 Hugging Face, 26 AI Hub, 6 GitHub. Examples: KSS Dataset, ClovaCall, KsponSpeech. 8 are marked commercial-use-allowed in the catalog. 8 rows look like ASR and 7 like TTS — confirm the original card before mixing them in one training run. Use this list to compare license and size, then open the original provider. GearDel does not host files.
Pick ASR for recognition, TTS for synthesis. Hours and speaker count sit in the size column — they are catalog strings, not remeasured audio.
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| KSS Dataset | Korean Speech Synthesis Dataset | Hugging Face | 12hours+ | CC BY-NC-SA 4.0 | Commercial use restricted |
| ClovaCall | Korean Speech Recognition Dataset | GitHub | 11,000 | MIT | Commercial use restricted |
| KsponSpeech | Korean Speech Recognition Dataset | AI Hub | 1000hours | AI Hub terms | Commercial use may be conditional |
| Speech Emotion Dataset (ko) | Korean Speech Dataset | GitHub | 3.2GB | License unknown | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 500hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 120hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 150hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 200hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub | 100hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 80hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 800hours | AI Hub terms | Commercial use may be conditional |
| TTS | Korean Speech Synthesis Dataset | AI Hub | 300hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 600hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 240GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 400hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub | 180hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 4TB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 50hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 40GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 15GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 30GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 85GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 120GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 60hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 150GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 8GB | AI Hub terms | Commercial use may be conditional |
| TTS | Korean Speech Synthesis Dataset | AI Hub | 90hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 220hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 140hours | AI Hub terms | Commercial use may be conditional |
| Zeroth | Korean Speech Recognition Dataset | GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| MeloTTS | Korean Speech Synthesis Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| KoSum | Korean Speech Dataset | Hugging Face (iontail) | 45GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| KoelLabs | Korean Speech Recognition Dataset | Hugging Face (KoelLabs) | 18 GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| IoT | Korean Speech Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub / GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub / GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Pansori TEDxKR | Korean Speech Recognition Dataset | GitHub | ~3hours | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| OLKAVS | Korean Speech Recognition Dataset | GitHub | 1,150hours | License unknown | Commercial use may be conditional |