Guide

Korean AI dataset licenses in this catalog

This page compares license labels that already appear on GearDel. It is not a law firm memo. Counts are from the live catalog (286 datasets).

Start with the commercial mark, then the license string

161 rows are marked allowed, 120 conditional, 5 restricted. “Allowed” on GearDel means the catalog inferred a permissive path from the recorded license — not that your product is cleared. A CC BY-SA dataset can still be commercial in some readings and painful for closed models in others.

182 license fields are verified against the original. 87 still need a human pass. Treat needs-review rows as a sticky note, not a green light.

Labels that show up most often

LabelRowsHow we read it here
MIT57Permissive in the usual MIT sense. Still open the original LICENSE file.
Apache 2.048Permissive with a patent grant. Confirm NOTICE files on the host.
CC BY-SA 4.026Share-alike. Derivatives often must stay under the same terms.
CC BY 4.012Attribution required. Not the same as CC BY-SA.
AI Hub terms92Not a Creative Commons license. Signup and AI Hub rules apply.
NIKL pledge7National Institute of Korean Language use pledge. Not OSI MIT.
Unlisted17Catalog does not claim a license. Do not assume it is open.

AI Hub is not “free on the internet”

92 rows carry the AI Hub terms label. Those pages usually require an account, a purpose, and a separate agreement. Putting them next to MIT GitHub repos in one table is the point of GearDel — and also the trap, if you download as if they were Apache 2.0.

NIKL pledge rows (7) are similar: the corpus is public as a research resource, not a drop-in open-source package.

A short working order

  1. Read the GearDel license field and its verified/needs-review badge.
  2. Open the original host. If the README disagrees, trust the host.
  3. Only then decide training vs evaluation. Benchmarks with ND clauses are easy to misuse as train sets.

FAQ → · How to choose a dataset → · Korean dataset collection →