Compare

Compare two catalog rows

The table copies license, commercial mark, provider, and size strings already stored for each dataset. It does not rank quality, and it does not remeasure files. Open the original page before you download.

See catalog-wide counts

APEACH · BEEP! Korean Hate Speech

APEACH is cataloged as Korean LLM Benchmark Dataset. BEEP! Korean Hate Speech is cataloged as Korean Toxicity Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Text.

License labels: APEACH CC BY-SA 4.0 (needs review); BEEP! Korean Hate Speech CC BY-SA 4.0 (verified).

Commercial-use marks: APEACH Commercial use allowed (catalog); BEEP! Korean Hate Speech Commercial use may be conditional. These marks are not legal advice.

Providers: APEACH Hugging Face (original URL verified); BEEP! Korean Hate Speech GitHub (original URL verified).

Size strings: APEACH 11,666 sentences, 1.1MB; BEEP! Korean Hate Speech 9.4K comments, 1.2MB. GearDel did not remeasure the files.

Year labels: APEACH 2022; BEEP! Korean Hate Speech 2020.

FieldAPEACHBEEP! Korean Hate Speech
TopicKorean LLM Benchmark DatasetKorean Toxicity Dataset
TypeTextText
LicenseCC BY-SA 4.0CC BY-SA 4.0
License checkneeds reviewverified
Commercial markCommercial use allowed (catalog)Commercial use may be conditional
ProviderHugging FaceGitHub
Original URLverifiedverified
LanguageKoreanKorean
FormatTextText
Size string11,666 sentences · 1.1MB9.4K comments · 1.2MB
Year20222020

Pairs already written out

These pairs stay in the HTML so the differences can be read without using the picker.

KMMLU · KoBEST

KMMLU is cataloged as Korean LLM Benchmark Dataset. KoBEST is cataloged as Korean LLM Benchmark Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Text.

License labels: KMMLU CC BY-ND 4.0 (verified); KoBEST CC BY-SA 4.0 (verified).

Commercial-use marks: KMMLU Commercial use restricted; KoBEST Commercial use allowed (catalog). These marks are not legal advice.

Providers: KMMLU Hugging Face (original URL verified); KoBEST Hugging Face (original URL verified).

Size strings: KMMLU 35,030 items, 70.7 MB; KoBEST 4.8K items, 5.1MB. GearDel did not remeasure the files.

Year labels: KMMLU 2024; KoBEST 2022.

FieldKMMLUKoBEST
TopicKorean LLM Benchmark DatasetKorean LLM Benchmark Dataset
TypeTextText
LicenseCC BY-ND 4.0CC BY-SA 4.0
License checkverifiedverified
Commercial markCommercial use restrictedCommercial use allowed (catalog)
ProviderHugging FaceHugging Face
Original URLverifiedverified
LanguageKoreanKorean
FormatTextText
Size string35,030 items · 70.7 MB4.8K items · 5.1MB
Year20242022

KMMLU · KMMLU-Pro

KMMLU is cataloged as Korean LLM Benchmark Dataset. KMMLU-Pro is cataloged as Korean LLM Benchmark Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Text.

License labels: KMMLU CC BY-ND 4.0 (verified); KMMLU-Pro CC BY-NC 4.0 (verified).

Commercial-use marks: KMMLU Commercial use restricted; KMMLU-Pro Commercial use allowed (catalog). These marks are not legal advice.

Providers: KMMLU Hugging Face (original URL verified); KMMLU-Pro Hugging Face (LGAI-EXAONE) (original URL verified).

Size strings: KMMLU 35,030 items, 70.7 MB; KMMLU-Pro 2,822 items, 12MB. GearDel did not remeasure the files.

Year labels: KMMLU 2024; KMMLU-Pro 2025.

FieldKMMLUKMMLU-Pro
TopicKorean LLM Benchmark DatasetKorean LLM Benchmark Dataset
TypeTextText
LicenseCC BY-ND 4.0CC BY-NC 4.0
License checkverifiedverified
Commercial markCommercial use restrictedCommercial use allowed (catalog)
ProviderHugging FaceHugging Face (LGAI-EXAONE)
Original URLverifiedverified
LanguageKoreanKorean
FormatTextText
Size string35,030 items · 70.7 MB2,822 items · 12MB
Year20242025

KsponSpeech · KSS Dataset

KsponSpeech is cataloged as Korean Speech Recognition Dataset. KSS Dataset is cataloged as Korean Speech Synthesis Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Speech.

License labels: KsponSpeech AI Hub terms (verified); KSS Dataset CC BY-NC-SA 4.0 (verified).

Commercial-use marks: KsponSpeech Commercial use may be conditional; KSS Dataset Commercial use restricted. These marks are not legal advice.

Providers: KsponSpeech AI Hub (original URL verified); KSS Dataset Hugging Face (original URL verified).

Size strings: KsponSpeech 2,000 · WAV/, 1000hours; KSS Dataset WAV 12,853, 12hours+. GearDel did not remeasure the files.

Year labels: KsponSpeech 2018; KSS Dataset 2018.

FieldKsponSpeechKSS Dataset
TopicKorean Speech Recognition DatasetKorean Speech Synthesis Dataset
TypeSpeechSpeech
LicenseAI Hub termsCC BY-NC-SA 4.0
License checkverifiedverified
Commercial markCommercial use may be conditionalCommercial use restricted
ProviderAI HubHugging Face
Original URLverifiedverified
LanguageKoreanKorean
FormatAudioAudio
Size string2,000 · WAV/ · 1000hoursWAV 12,853 · 12hours+
Year20182018

KLUE Benchmark · KoBEST

KLUE Benchmark is cataloged as Korean LLM Benchmark Dataset. KoBEST is cataloged as Korean LLM Benchmark Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Text.

License labels: KLUE Benchmark CC BY-SA 4.0 (verified); KoBEST CC BY-SA 4.0 (verified).

Commercial-use marks: KLUE Benchmark Commercial use may be conditional; KoBEST Commercial use allowed (catalog). These marks are not legal advice.

Providers: KLUE Benchmark Hugging Face (original URL verified); KoBEST Hugging Face (original URL verified).

Size strings: KLUE Benchmark 8, 62.3 MB; KoBEST 4.8K items, 5.1MB. GearDel did not remeasure the files.

Year labels: KLUE Benchmark 2021; KoBEST 2022.

FieldKLUE BenchmarkKoBEST
TopicKorean LLM Benchmark DatasetKorean LLM Benchmark Dataset
TypeTextText
LicenseCC BY-SA 4.0CC BY-SA 4.0
License checkverifiedverified
Commercial markCommercial use may be conditionalCommercial use allowed (catalog)
ProviderHugging FaceHugging Face
Original URLverifiedverified
LanguageKoreanKorean
FormatTextText
Size string8 · 62.3 MB4.8K items · 5.1MB
Year20212022

KoAlpaca-v1.1a · KULLM-v2

KoAlpaca-v1.1a is cataloged as Korean Instruction Tuning Dataset. KULLM-v2 is cataloged as Korean Instruction Tuning Dataset. Topics are inferred from the title and tags, so they can differ from the host’s official task name.

Both rows are type Text.

License labels: KoAlpaca-v1.1a License unknown (unknown); KULLM-v2 Apache 2.0 (verified).

Commercial-use marks: KoAlpaca-v1.1a Commercial use allowed (catalog); KULLM-v2 Commercial use allowed (catalog). These marks are not legal advice.

Providers: KoAlpaca-v1.1a Hugging Face (original URL verified); KULLM-v2 Hugging Face (original URL verified).

Size strings: KoAlpaca-v1.1a 21,155, Large; KULLM-v2 152,630, 180MB. GearDel did not remeasure the files.

Year labels: KoAlpaca-v1.1a 2023; KULLM-v2 2023.

FieldKoAlpaca-v1.1aKULLM-v2
TopicKorean Instruction Tuning DatasetKorean Instruction Tuning Dataset
TypeTextText
LicenseLicense unknownApache 2.0
License checkunknownverified
Commercial markCommercial use allowed (catalog)Commercial use allowed (catalog)
ProviderHugging FaceHugging Face
Original URLverifiedverified
LanguageKoreanKorean
FormatTextText
Size string21,155 · Large152,630 · 180MB
Year20232023