Collection

Korean speech recognition (ASR) datasets

8 catalog items · GearDel

GearDel lists 8 catalog items in this group (8 speech). Sources: 3 AI Hub, 4 GitHub. Examples: ClovaCall, KsponSpeech, Korean Speech Recognition Dataset. 2 are marked commercial-use-allowed in the catalog. These rows look like recognition (ASR), not read-speech TTS. Hours are catalog strings, not remeasured audio. Use this list to compare license and size, then open the original provider. GearDel does not host files.

Unlike the full speech collection, this table keeps rows that look like recognition (ASR). Spontaneous sets such as KsponSpeech can sit next to read-speech ASR such as Zeroth. Do not mix them with TTS read-speech in one run.

Hours and speaker counts are catalog strings. Hosts sometimes publish samples or split archives — do not treat the hour cell as local wav length.

Call-center and dialect audio often have consent rules stricter than the license line. A commercial-allowed mark still needs the host’s speaker-consent section.

If you need TTS, go back to the speech collection and keep synthesis rows. GearDel does not host files.

Recognition only. If you need TTS, use the speech collection and keep rows whose topic says synthesis.

Allowed

2

Conditional

5

Restricted

1

Counts are rows in this collection, not downloads. License labels below are the stored strings.

NameUse / topicSourceSizeLicenseCommercial
ClovaCallKorean Speech Recognition DatasetGitHub11,000MITCommercial use restricted
KsponSpeechKorean Speech Recognition DatasetAI Hub1000hoursAI Hub termsCommercial use may be conditional
Korean Speech Recognition DatasetKorean Speech Recognition DatasetAI Hub120hoursAI Hub termsCommercial use may be conditional
Korean Speech Recognition DatasetKorean Speech Recognition DatasetAI Hub150hoursAI Hub termsCommercial use may be conditional
ZerothKorean Speech Recognition DatasetGitHubLargeApache 2.0Commercial use allowed (catalog)
KoelLabsKorean Speech Recognition DatasetHugging Face (KoelLabs)18 GBCC BY-NC 4.0Commercial use allowed (catalog)
Pansori TEDxKRKorean Speech Recognition DatasetGitHub~3hoursCC BY-NC-ND 4.0Commercial use may be conditional
OLKAVSKorean Speech Recognition DatasetGitHub1,150hoursLicense unknownCommercial use may be conditional

Other collections