Guide
How to pick a Korean AI dataset on GearDel
The catalog has 286 rows. The failure mode is not “too few datasets”. It is downloading a benchmark as a train set, or an AI Hub dump as if it were MIT.
1. Name the task before you name the dataset
GearDel types in this snapshot: text 215, image 33, speech 38. OCR is not “any image set”. ASR is not TTS. Instruction tuning is not pretraining. If you cannot say which of those you need, the search box will happily return famous names that do not fit.
Use collections when the task is clear: speech, OCR, instruction, benchmarks. Use the homepage filters when you already know a license string or a host.
2. Split evaluation from training
KLUE, KoBEST, and KMMLU-family pages in this catalog are scoring sets. Training on the test split makes a leaderboard number, not a model. If the collection title says benchmark, do not treat the table as a pretraining mix.
3. Read license before size
A 1,000-hour speech set you cannot use commercially is not a bargain. 161 rows are marked allowed; 92 still sit under AI Hub terms. Size strings on GearDel are catalog text copied from the host, not a re-measure.
4. Trust the original link
130 source URLs are verified. If the badge is needs-review or unknown, search the provider by dataset name instead of guessing a Hugging Face id. GearDel never hosts the archive; a 404 on our side after a host deletion should 301 home, not keep a dead card.
5. A 10-minute path that usually works
- Open a collection that matches the task.
- Sort mentally: verified license first, then source you can actually access (HF vs AI Hub account).
- Open two or three dataset pages, not twenty. Read the “when this fits” paragraph.
- Leave GearDel for the original README. If that page disagrees, the README wins.