← Open Source
taishi-i

awesome-japanese-nlp-resources

A curated list of resources for Japanese natural language processing (NLP): Python libraries, LLMs, dictionaries, corpora, and datasets. Includes Claude Code and Codex skills to search resources.

ListsTool collectionsModel collections
Open on GitHub
Momentum
+1stars in 24 hours+0.1%
1.02k
Stars
54
Forks
+7
This week
12
Contributors
Created 2022-06-08 · Updated 2026-10-07 · #5260 today
Top developers
README

awesome-japanese-nlp-resources

Awesome RRs License: CC0-1.0 CC0

A curated list of resources dedicated to Python libraries, llms, dictionaries, and corpora of NLP for Japanese

English | 日本語 (Japanese) | 繁體中文 (Chinese) | 简体中文 (Chinese)

🎉 The latest additions

Python

  • yomiyasu - AI生成の日本語を自然な日本語へ推敲するAgent Skill / Agent Skill for Refining AI-Generated Japanese into Natural Japanese

Dictionary and IME

  • Meltype - 半角/全角 キーを押さなくても、日本語と英語を打ち分けられるようにする Windows 常駐ツールです。

Updated on Oct 06, 2026

Claude Code & Codex Plugin

Search, discover, and track Japanese NLP resources directly from Claude Code or Codex using the awesome-japanese-nlp-resources plugin. Both tools use the same skills and the same bundled data.

Claude Code:

# Add the marketplace and install
/plugin marketplace add taishi-i/awesome-japanese-nlp-resources
/plugin install awesome-japanese-nlp-resources@awesome-japanese-nlp-resources
/reload-plugins

Codex (run in your terminal, then start a new Codex session):

codex plugin marketplace add taishi-i/awesome-japanese-nlp-resources
codex plugin add awesome-japanese-nlp-resources@awesome-japanese-nlp-resources

The plugin ships four skills:

Skill Purpose
/awesome-japanese-nlp-resources:search Search the bundled 1,200+ resource dataset
/awesome-japanese-nlp-resources:discover Given a tool, find alternatives (listed and unlisted); given a topic, find contribution candidates not yet in the list
/awesome-japanese-nlp-resources:compare Compare several libraries/models/datasets across a few criteria as a ○/△/✕ table
/awesome-japanese-nlp-resources:research Survey the dataset + latest web research for a combined trend + challenges report

In Codex, call the same skills with $ instead of / (e.g. $awesome-japanese-nlp-resources:search ).

All skills detect the query language and respond in kind — English by default, Japanese when the query contains Japanese characters.

search example:

/awesome-japanese-nlp-resources:search What tutorials are recommended for beginners learning Japanese NLP?

Output includes a Use-case Selection Guide table:

Use case Recommended Popularity Why
Learn modern NLP/LLM concepts from scratch llm-book ⭐387 Most popular intro; covers LLMs end-to-end with practical Japanese code examples
Build practical skills through exercises nlp100v2025 ⭐82 Classic "100 Knock" challenge series; best way to learn by doing
Learn BERT and transformer-based NLP bert-book ⭐262 Hands-on BERT programming with clear explanations for Japanese NLP beginners

For full documentation, see the plugin README.

Contents

Python library

Morphology analysis

Libraries that split Japanese text into words or morphemes and assign part-of-speech and base forms

  • sudachi.rs - SudachiPy 0.6* and above are developed as Sudachi.rs.
  • Janome - Japanese morphological analysis engine written in pure Python
  • mecab-python3 - mecab-python. you can find original version here:http://taku910.github.io/mecab/
  • mecab - This repository is for building Windows 64-bit MeCab binary and improving MeCab Python binding.
  • fugashi - A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.
  • nagisa - A Japanese tokenizer based on recurrent neural networks
  • pyknp - A Python Module for JUMAN++/KNP
  • Mykytea-python - Python wrapper for KyTea
  • konoha - Konoha: Simple wrapper of Japanese Tokenizers
  • natto-py - natto-py combines the Python programming language with MeCab, the part-of-speech and morphological analyzer for the Japanese language.
  • rakutenma-python - Rakuten MA (Python version)
  • python-vaporetto - Vaporetto is a fast and lightweight pointwise prediction based tokenizer. This is a Python wrapper for Vaporetto.
  • dango - An easy to use tokenizer for Japanese text, aimed at language learners and non-linguists
  • rhoknp - Yet another Python binding for Juman++/KNP
  • python-vibrato - Viterbi-based accelerated tokenizer (Python wrapper)
  • jagger-python - Python binding for Jagger(C++ implementation of Pattern-based Japanese Morphological Analyzer)
  • Mecari - Mecari (Japanese Morphological Analysis with Graph Neural Networks)
Name downloads/week total downloads stars last commit
🔗 SudachiPy 📥 332k 📦 75M ⭐ 443 🔴 october 2022
🔗 Janome 📥 50k 📦 13M ⭐ 917 🟡 june
🔗 mecab-python3 📥 149k 📦 42M ⭐ 579 🟡 november 2025
🔗 mecab 📥 91k 📦 3M ⭐ 274 🔴 october 2024
🔗 fugashi 📥 185k 📦 18M ⭐ 536 🟡 october 2025
🔗 nagisa 📥 42k 📦 10M ⭐ 420 🟢 september
🔗 pyknp 📥 616 📦 3M ⭐ 93 🟡 january
🔗 Mykytea-python 📥 904 📦 590k ⭐ 36 🟡 march
🔗 konoha 📥 35k 📦 7M ⭐ 265 🟢 september
🔗 natto-py 📥 1k 📦 34M ⭐ 95 🔴 november 2023
🔗 rakutenma-python 📥 30 📦 27k ⭐ 23 🔴 may 2017
🔗 python-vaporetto 📥 669 📦 191k ⭐ 21 🟡 may
🔗 dango 📥 27 📦 27k ⭐ 26 🔴 november 2021
🔗 rhoknp 📥 13k 📦 2M ⭐ 40 🟢 september
🔗 python-vibrato 📥 269 📦 128k ⭐ 46 🟡 may
🔗 jagger-python 📥 1k 📦 337k ⭐ 13 🔴 march 2024
🔗 Mecari - - ⭐ 41 🔴 september 2025

Parsing

Libraries that analyze syntactic and dependency structures of Japanese sentences

  • ginza - A Japanese NLP Library using spaCy as framework based on Universal Dependencies
  • cabocha - Yet Another Japanese Dependency Structure Analyzer
  • UniDic2UD - Tokenizer POS-tagger Lemmatizer and Dependency-parser for modern and contemporary Japanese
  • camphr - Camphr - NLP libary for creating pipeline components
  • SuPar-UniDic - Tokenizer POS-tagger Lemmatizer and Dependency-parser for modern and contemporary Japanese with BERT models
  • depccg - A* CCG Parser with a Supertag and Dependency Factored Model
  • bertknp - A Japanese dependency parser based on BERT
  • esupar - Tokenizer POS-Tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models for Japanese and other languages
  • yomikata - Heteronym disambiguation library using a fine-tuned BERT model.
  • jdepp-python - Python binding for J.DepP(C++ implementation of Japanese Dependency Parsers)
  • lightblue - A CCG parser for Japanese with DTS-representations
  • natsume-simple - natsume-simpleは日本語の係り受け関係検索システム
  • jdeppy - Python wrapper for J.DepP, fast Japanese Dependency Parser
Name downloads/week total downloads stars last commit
🔗 ginza 📥 12k 📦 2M ⭐ 873 🟢 last wednesday
🔗 cabocha 📥 64 📦 56k ⭐ 7 🔴 august 2022
🔗 UniDic2UD 📥 337 📦 340k ⭐ 38 🟡 december 2025
🔗 camphr 📥 433 📦 280k ⭐ 336 🔴 august 2021
🔗 SuPar-UniDic 📥 118 📦 125k ⭐ 21 🟢 august
🔗 depccg 📥 162 📦 52k ⭐ 103 🟢 august
🔗 bertknp - - ⭐ 23 🔴 october 2021
🔗 esupar 📥 411 📦 186k ⭐ 55 🟢 september
🔗 yomikata 📥 30 📦 51k ⭐ 36 🔴 october 2023
🔗 jdepp-python 📥 2k 📦 320k ⭐ 4 🔴 february 2024
🔗 lightblue - - ⭐ 29 🟡 march
🔗 natsume-simple - - ⭐ 5 🔴 february 2025
🔗 jdeppy 📥 24 📦 12k ⭐ 3 🔴 february 2022

Converter

Libraries that convert between character types such as kana, romaji, and full-width/half-width forms

  • pykakasi - Lightweight converter from Japanese Kana-kanji sentences into Kana-Roman.
  • cutlet - Japanese to romaji converter in Python
  • alphabet2kana - Convert English alphabet to Katakana
  • Convert-Numbers-to-Japanese - Converts Arabic numerals, or 'western' style numbers, to a Japanese context.
  • mozcpy - Mozc for Python: Kana-Kanji converter
  • jamorasep - Japanese text parser to separate Hiragana/Katakana string into morae (syllables).
  • text2phoneme - 日本語文を音素列へ変換するスクリプト
  • jntajis-python - A fast character conversion and transliteration library based on the scheme defined for Japan National Tax Agency (国税庁) 's
  • wiredify - Convert japanese kana from ba-bi-bu-be-bo into va-vi-vu-ve-vo
  • mecab-text-cleaner - Simple Python package (CLI/Python API) for getting japanese readings (yomigana) and accents using MeCab.
  • pynormalizenumexp - 数量表現や時間表現の抽出・正規化を行うNormalizeNumexpのPython実装
  • Jusho - Easy wrapper for the postal code data of Japan
  • yurenizer - Japanese text normalizer that resolves spelling inconsistencies. (日本語表記揺れ解消ツール)
  • e2k - A tool for automatic English to Katakana conversion
  • alkana.py - A tool to get the katakana reading of an alphabetical string.
  • englishtokanaconverter - 英語文字列をカタカナに変換するプログラム
  • kanjiconv - Kanji Converter to Hiragana, Katakana, Roman alphabet.
  • kanjize - Kanjize(カンジャイズ): Easy converter between Kanji-Number and Integer
Name downloads/week total downloads stars last commit
🔗 pykakasi 📥 565k 📦 40M ⭐ invalid 🟢 july
🔗 cutlet 📥 20k 📦 2M ⭐ 387 🟢 september
🔗 alphabet2kana 📥 168 📦 65k ⭐ 15 🟡 february
🔗 Convert-Numbers-to-Japanese - - ⭐ 51 🔴 november 2020
🔗 mozcpy 📥 249 📦 21k ⭐ 49 🔴 february 2025
🔗 jamorasep 📥 110 📦 12k ⭐ 11 🟡 february
🔗 text2phoneme - - ⭐ 13 🔴 may 2023
🔗 jntajis-python 📥 1k 📦 148k ⭐ 21 🟡 march
🔗 wiredify 📥 28 📦 7k ⭐ 3 🟡 december 2025
🔗 mecab-text-cleaner 📥 33 📦 5k ⭐ 7 🟢 september
🔗 pynormalizenumexp 📥 53 📦 15k ⭐ 8 🔴 april 2024
🔗 Jusho 📥 360 📦 69k ⭐ 12 🔴 june 2024
🔗 yurenizer 📥 96 📦 22k ⭐ 6 🔴 march 2025
🔗 e2k 📥 828 📦 42k ⭐ 20 🟡 march
🔗 alkana.py - - ⭐ 35 🔴 october 2021
🔗 englishtokanaconverter - - ⭐ 4 🟢 august
🔗 kanjiconv 📥 660 📦 29k ⭐ 20 🟢 august
🔗 kanjize 📥 15k 📦 2M ⭐ 68 🔴 june 2025

Preprocessor

Libraries that normalize and clean text before analysis

  • neologdn - Japanese text normalizer for mecab-neologd
  • jaconv - Pure-Python Japanese character interconverter for Hiragana, Katakana, Hankaku, and Zenkaku
  • mojimoji - A fast converter between Japanese hankaku and zenkaku characters
  • text-cleaning - A powerful text cleaner for Japanese web texts
  • HojiChar - 複数の前処理を構成して管理するテキスト前処理ツール
  • utsuho - Utsuho is a Python module that facilitates bidirectional conversion between half-width katakana and full-width katakana in Japanese.
  • python-habachen - Yet Another Fast Japanese String Converter
  • kairyou - Quickly preprocesses Japanese text using NLP/NER from SpaCy for Japanese translation or other NLP tasks.
Name downloads/week total downloads stars last commit
🔗 neologdn 📥 11k 📦 2M ⭐ 290 🟡 may
🔗 jaconv 📥 852k 📦 84M ⭐ 351 🟡 february
🔗 mojimoji 📥 60k 📦 13M ⭐ 154 🔴 january 2024
🔗 text-cleaning - - ⭐ 12 🔴 november 2022
🔗 HojiChar 📥 833 📦 1M ⭐ 127 🟢 september
🔗 utsuho 📥 83 📦 26k ⭐ 5 🟡 march
🔗 python-habachen 📥 18k 📦 3M ⭐ 6 🟡 october 2025
🔗 kairyou 📥 158 📦 35k ⭐ 6 🔴 june 2025

Sentence splitter

Libraries that automatically detect sentence boundaries and split text

  • Bunkai - Sentence boundary disambiguation tool for Japanese texts (日本語文境界判定器)
  • japanese-sentence-breaker - Japanese Sentence Breaker
  • sengiri - Yet another sentence-level tokenizer for the Japanese text
  • budoux - Standalone. Small. Language-neutral. BudouX is the successor to Budou, the machine learning powered line break organizer tool.
  • ja_sentence_segmenter - japanese sentence segmentation library for python
  • hasami - A tool to perform sentence segmentation on Japanese text
  • kuzukiri - Japanese Text Segmenter for Python written in Rust
  • ja-senter-benchmark - Comparison of Japanese Sentence Segmentation Tools
  • fast-bunkai - Japanese sentence splitting(日本語文境界判定器), 40–250× faster via a Rust-accelerated Python library with near-perfect API compatibility with megagonlabs/bunkai.
Name downloads/week total downloads stars last commit
🔗 bunkai 📥 340 📦 124k ⭐ 200 🔴 august 2023
🔗 japanese-sentence-breaker 📥 16 📦 7k ⭐ 14 🔴 february 2021
🔗 sengiri 📥 197 📦 139k ⭐ 24 🟡 november 2025
🔗 budoux 📥 17k 📦 820k ⭐ 1.8k 🟢 today
🔗 ja_sentence_segmenter 📥 3k 📦 281k ⭐ 76 🟢 july
🔗 hasami 📥 103 📦 47k ⭐ 6 🔴 february 2021
🔗 kuzukiri 📥 107 📦 31k ⭐ 6 🔴 june 2025
🔗 ja-senter-benchmark - - ⭐ 10 🔴 february 2023
🔗 fast-bunkai 📥 101 📦 10k ⭐ 77 🟡 october 2025

Sentiment analysis

Libraries that detect emotions or polarity in text

  • oseti - Dictionary based Sentiment Analysis for Japanese
  • negapoji - Japanese negative positive classification.日本語文書のネガポジを判定。
  • pymlask - Emotion analyzer for Japanese text
  • asari - Japanese sentiment analyzer implemented in Python.
  • kotobacore - Japanese semantic understanding engine for LLM, RAG, SNS analysis, and AI agents ? dependency-free tokenizer + Plutchik emotion + intent + RAG keywords.
Name downloads/week total downloads stars last commit
🔗 oseti 📥 273 📦 175k ⭐ 99 🔴 august 2025
🔗 negapoji - - ⭐ 152 🔴 august 2017
🔗 pymlask 📥 52 📦 68k ⭐ 119 🔴 july 2024
🔗 asari 📥 78 📦 83k ⭐ 153 🔴 october 2022
🔗 kotobacore - - ⭐ 1 🟢 last tuesday

Machine translation

Libraries that automatically translate text between languages

  • jparacrawl-finetune - An example usage of JParaCrawl pre-trained Neural Machine Translation (NMT) models.
  • JASS - JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation (LREC2020) & Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation (ACM TALLIP)
  • PheMT - A phenomenon-wise evaluation dataset for Japanese-English machine translation robustness. The dataset is based on the MTNT dataset, with additional annotations of four linguistic phenomena; Proper Noun, Abbreviated Noun, Colloquial Expression, and Variant. COLING 2020.
  • VISA - An ambiguous subtitles dataset for visual scene-aware machine translation
  • plamo-translate-cli - A command-line interface for translation using the plamo-2-translate model with local execution.
Name downloads/week total downloads stars last commit
🔗 jparacrawl-finetune - - ⭐ 104 🔴 april 2021
🔗 JASS - - ⭐ 16 🔴 january 2022
🔗 PheMT - - ⭐ 20 🔴 february 2021
🔗 VISA - - ⭐ 14 🔴 october 2022
🔗 plamo-translate-cli - - ⭐ 351 🟡 april

Named entity recognition

Libraries that extract names of people, places, and organizations from text

  • namaco - Character Based Named Entity Recognition.
  • entitypedia - Entitypedia is an Extended Named Entity Dictionary from Wikipedia.
  • noyaki - Converts character span label information to tokenized text-based label information.
  • bert-japanese-ner-finetuning - Code to perform finetuning of the BERT model. BERTモデルのファインチューニングで固有表現抽出用タスクのモデルを作成・使用するサンプルです
  • joint-information-extraction-hs - 詳細なアノテーション基準に基づく症例報告コーパスからの固有表現及び関係の抽出精度の推論を行うコード
  • pygeonlp - pygeonlp, A python module for geotagging Japanese texts.
  • bert-ner-japanese - BERTによる日本語固有表現抽出のファインチューニング用プログラム
  • huggingface-finetune-japanese - Examples to finetune encoder-only and encoder-decoder transformers for Japanese language (Hugging Face) Resources
  • novelanalysisbyner - BERTのfine-tuningによる固有表現抽出
Name downloads/week total downloads stars last commit
🔗 namaco - - ⭐ 40 🔴 february 2018
🔗 entitypedia - - ⭐ 13 🔴 december 2018
🔗 noyaki 📥 179 📦 25k ⭐ 5 🔴 august 2022
🔗 bert-japanese-ner-finetuning - - ⭐ 11 🔴 june 2022
🔗 joint-information-extraction-hs - - ⭐ 2 🔴 november 2021
🔗 pygeonlp 📥 367 📦 30k ⭐ 22 🟡 march
🔗 bert-ner-japanese - - ⭐ 5 🔴 september 2022
🔗 huggingface-finetune-japanese - - ⭐ 16 🔴 october 2023
🔗 novelanalysisbyner - - ⭐ 2 🔴 june 2024

OCR

Libraries that recognize and extract text from images

  • Manga OCR - About Optical character recognition for Japanese text, with the main focus being Japanese manga
  • mokuro - Read Japanese manga inside browser with selectable text.
  • handwritten-japanese-ocr - Handwritten Japanese OCR demo using touch panel to draw the input text using Intel OpenVINO toolkit
  • OCR_Japanease - 日本語OCR
  • ndlocr_cli - NDLOCRのアプリケーション
  • donut - Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
  • JMTrans - manga translator - get japanese manga from url to translate manga image
  • Kindai-OCR - OCR system for recognizing modern Japanese magazines
  • text_recognition - NDLOCR用テキスト認識モジュール
  • Poricom - Optical character recognition in manga images. Manga OCR desktop application
  • owocr - Optical character recognition for Japanese text
  • yomitoku - Yomitoku is an AI-powered document image analysis package designed specifically for the Japanese language.
  • findtextcenternet - Japanese OCR with CenterNet
  • simple-ocr-for-manga - A simple OCR for manga (Japanese traditional and Japanese vertical)
  • jp-ocr-evaluation - 日本語の文章画像に対するOCRの性能を評価
  • paddleocr-vl-sft-for-japanese-manga-on-rtx-3060 - Fine-tune PaddleOCR-VL on the Manga109s dataset for Japanese manga text recognition. The base model struggles with vertical Japanese text reading order in manga. After fine-tuning, the model correctly handles manga-specific text layouts.
  • MangaOCR - A lightweight OCR model for Japanese text, especially in Manga
  • meikiocr - high-speed, high-accuracy, local ocr for japanese video games
  • meikipop - universal japanese ocr popup dictionary for windows, linux and macos
Name downloads/week total downloads stars last commit
🔗 manga-ocr 📥 3k 📦 398k ⭐ 2.8k 🟢 july
🔗 mokuro 📥 464 📦 114k ⭐ 1.7k 🟢 july
🔗 handwritten-japanese-ocr - - ⭐ 37 🔴 april 2022
🔗 OCR_Japanease - - ⭐ 251 🔴 april 2021
🔗 ndlocr_cli - - ⭐ 682 🔴 september 2025
🔗 donut 📥 98 📦 205k ⭐ 6.9k 🔴 july 2023
🔗 JMTrans - - ⭐ 87 🔴 january 2021
🔗 Kindai-OCR - - ⭐ 153 🟢 july
🔗 text_recognition - - ⭐ 8 🔴 july 2023
🔗 Poricom - - ⭐ 440 🔴 june 2023
🔗 owocr - - ⭐ 301 🟡 march
🔗 yomitoku 📥 1k 📦 133k ⭐ 1.6k 🟢 last thursday
🔗 findtextcenternet - - ⭐ 65 🔴 august 2025
🔗 simple-ocr-for-manga - - ⭐ 7 🔴 march 2025
🔗 jp-ocr-evaluation - - ⭐ 1 🔴 march 2024
🔗 paddleocr-vl-sft-for-japanese-manga-on-rtx-3060 - - ⭐ 15 🟡 december 2025
🔗 MangaOCR - - ⭐ 40 🔴 may 2024
🔗 meikiocr 📥 801 📦 54k ⭐ 98 🟢 september
🔗 meikipop - - ⭐ 679 🟢 september

Tool for pretrained models

Libraries that utilize pretrained models to improve accuracy and efficiency

Name downloads/week total downloads stars last commit
🔗 JGLUE - - ⭐ 349 🔴 march 2025
🔗 ginza-transformers 📥 532 📦 279k ⭐ 16 🟢 last wednesday
🔗 t5_japanese_dialogue_generation - - ⭐ 3 🔴 november 2021
🔗 japanese_text_classification - - ⭐ 9 🔴 january 2020
🔗 Japanese-BERT-Sentiment-Analyzer - - ⭐ 2 🔴 april 2021
🔗 jmlm_scoring - - ⭐ 5 🔴 february 2022
🔗 allennlp-shiba-model 📥 39 📦 23k ⭐ 12 🔴 june 2021
🔗 evaluate_japanese_w2v - - ⭐ 12 🔴 november 2024
🔗 gector-ja - - ⭐ 20 🔴 june 2021
🔗 Japanese-BPEEncoder - - ⭐ 42 🔴 september 2021
🔗 Japanese-BPEEncoder_V2 - - ⭐ 41 🔴 january 2023
🔗 transformer-copy - - ⭐ 29 🔴 september 2020
🔗 nagisa_bert 📥 28 📦 63k ⭐ 5 🟡 february
🔗 JGLUE-benchmark - - ⭐ 18 🟢 last thursday
🔗 jptranstokenizer 📥 127 📦 31k ⭐ 5 🔴 february 2024
🔗 jp-stable - - ⭐ 155 🔴 november 2023
🔗 compare-ja-tokenizer - - ⭐ 6 🔴 june 2023
🔗 lm-evaluation-harness-jp-stable - - ⭐ 1 🔴 june 2023
🔗 llm-lora-classification - - ⭐ 97 🔴 july 2023
🔗 rinna_gpt-neox_ggml-lora - - ⭐ 19 🔴 may 2023
🔗 japanese-llm-roleplay-benchmark - - ⭐ 42 🔴 november 2023
🔗 japanese-llm-ranking - - ⭐ 50 🔴 march 2024
🔗 llm-jp-eval - - ⭐ 172 🟢 august
🔗 llm-jp-sft - - ⭐ 64 🔴 june 2024
🔗 llm-jp-tokenizer - - ⭐ 52 🟢 september
🔗 japanese-lm-fin-harness - - ⭐ 80 🟡 june
🔗 ja-vicuna-qa-benchmark - - ⭐ 33 🔴 june 2024
🔗 swallow-evaluation - - ⭐ 26 🔴 september 2025
🔗 swallow-evaluation-instruct - - ⭐ 34 🟡 may
🔗 pretrained_doc2vec_ja - - ⭐ 25 🔴 january 2019
🔗 pl-bert-ja - - ⭐ 24 🔴 december 2023

Others

General-purpose tools supporting Japanese language processing

Show 212 items

  • namedivider-python - A tool for dividing the Japanese full name into a family name and a given name.
  • asa-python - A curated list of resources dedicated to Python libraries of NLP for Japanese
  • python_asa - python版日本語意味役割付与システム(ASA)
  • toiro - A comparison tool of Japanese tokenizers
  • ja-timex - 自然言語で書かれた時間情報表現を抽出/規格化するルールベースの解析器
  • JapaneseTokenizers - A set of metrics for feature selection from text data
  • daaja - This repository has implementations of data augmentation for NLP for Japanese.
  • accel-brain-code - The purpose of this repository is to make prototypes as case study in the context of proof of concept(PoC) and research and development(R&D) that I have written in my website. The main research topics are Auto-Encoders in relation to the representation learning, the statistical machine learning for energy-based models, adversarial generation net…
  • kyoto-reader - A processor for KyotoCorpus, KWDLC, and AnnotatedFKCCorpus
  • nlplot - Visualization Module for Natural Language Processing
  • rake-ja - Rapid Automatic Keyword Extraction algorithm for Japanese
  • jel - Japanese Entity Linker.
  • MedNER-J - Latest version of MedEX/J (Japanese disease name extractor)
  • zunda-python - Zunda: Japanese Enhanced Modality Analyzer client for Python.
  • AIO2_DPR_baseline - https://www.nlp.ecei.tohoku.ac.jp/projects/aio/
  • showcase - A PyTorch implementation of the Japanese Predicate-Argument Structure (PAS) analyser presented in the paper of Matsubayashi & Inui (2018) with some improvements.
  • darts-clone-python - Darts-clone python binding
  • jrte-corpus_example - Example codes for Japanese Realistic Textual Entailment Corpus
  • desuwa - Feature annotator to morphemes and phrases based on KNP rule files (pure-Python)
  • HotPepperGourmetDialogue - Restaurant Search System through Dialogue in Japanese.
  • nlp-recipes-ja - Samples codes for natural language processing in Japanese
  • Japanese_nlp_scripts - Small example scripts for working with Japanese texts in Python
  • DNorm-J - Japanese version of DNorm
  • pyknp-eventgraph - EventGraph is a development platform for high-level NLP applications in Japanese.
  • ishi - Ishi: A volition classifier for Japanese
  • python-npylm - ベイズ階層言語モデルによる教師なし形態素解析
  • python-npycrf - 条件付確率場とベイズ階層言語モデルの統合による半教師あり形態素解析
  • unsupervised-pos-tagging - 教師なし品詞タグ推定
  • negima - Negima is a Python package to extract phrases in Japanese text by using the part-of-speeches based rules you defined.
  • YouyakuMan - Extractive summarizer using BertSum as summarization model
  • japanese-numbers-python - A parser for Japanese number (Kanji, arabic) in the natural language.
  • kantan - Lookup japanese words by radical patterns
  • make-meidai-dialogue - Get Japanese dialogue corpus
  • japanese_summarizer - A summarizer for Japanese articles.
  • chirptext - ChirpText is a collection of text processing tools for Python.
  • yubin - Japanese Address Munger
  • jawiki-cleaner - Japanese Wikipedia Cleaner
  • japanese2phoneme - A python library to convert Japanese to phoneme.
  • anlp_nlp2021_d3-1 - This repository contains codes related to the experiments in "An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification"
  • aozora_classification - This project aims to classify Japanese sentence to how well similar to some Japanese classical writers, such as Soseki Natsume, Ogai Mori, Ryunosuke Akutagawa and so on.
  • aozora-corpus-generator - Generates plain or tokenized text files from the Aozora Bunko
  • JLM - A fast LSTM Language Model for large vocabulary language like Japanese and Chinese
  • NTM - Testing of Neural Topic Modeling for Japanese articles
  • EN-JP-ML-Lexicon - This is a English-Japanese lexicon for Machine Learning and Deep Learning terminology.
  • text-generation - Easy-to-use scripts to fine-tune GPT-2-JA with your own texts, to generate sentences, and to tweet them automatically.
  • chainer_nic - Neural Image Caption (NIC) on chainer, its pretrained models on English and Japanese image caption datasets.
  • unihan-lm - The official repository for "UnihanLM: Coarse-to-Fine Chinese-Japanese Language Model Pretraining with the Unihan Database", AACL-IJCNLP 2020
  • mbart-finetuning - Code to perform finetuning of the mBART model.
  • xvector_jtubespeech - xvector model on jtubespeech
  • TinySegmenterMaker - TinySegmenter用の学習モデルを自作するためのツール.
  • Grongish - 日本語とグロンギ語の相互変換スクリプト
  • WordCloud-Japanese - WordCloudでの日本語文章をMecab(形態素解析エンジン)を使用せずに形態素解析チックな表示を実現するスクリプト
  • snark - 日本語ワードネットを利用したDBアクセスライブラリ
  • toEmoji - 日本語文を絵文字だけの文に変換するなにか
  • termextract - - 専門用語抽出アルゴリズムの実装の練習
  • JDT-with-KenLM-scoring - Japanese-Dialog-Transformerの応答候補に対して、KenLMによるN-gram言語モデルでスコアリングし、フィルタリング若しくはリランキングを行う。
  • mixture-of-unigram-model - Mixture of Unigram Model and Infinite Mixture of Unigram Model in Python. (混合ユニグラムモデルと無限混合ユニグラムモデル)
  • hidden-markov-model - Hidden Markov Model (HMM) and Infinite Hidden Markov Model (iHMM) in Python. (隠れマルコフモデルと無限隠れマルコフモデル)
  • Ngram-language-model - Ngram language model in Python. (Nグラム言語モデル)
  • ASRDeepSpeech - Automatic Speech Recognition with deepspeech2 model in pytorch with support from Zakuro AI.
  • neural_ime - Neural IME: Neural Input Method Engine
  • neural_japanese_transliterator - Can neural networks transliterate Romaji into Japanese correctly?
  • tinysegmenter - tokenizer specified for Japanese
  • AugLy-jp - Data Augmentation for Japanese Text on AugLy
  • furigana4epub - A Python script for adding furigana to Japanese epub books using Mecab and Unidic.
  • PyKatsuyou - Japanese verb/adjective inflections tool
  • jageocoder - Pure Python Japanese address geocoder
  • nksnd - New kana-kanji conversion engine
  • JaMIE - A Japanese Medical Information Extraction Toolkit
  • fasttext-vs-word2vec-on-twitter-data - fasttextとword2vecの比較と、実行スクリプト、学習スクリプトです
  • minimal-search-engine - 最小のサーチエンジン/PageRank/tf-idf
  • 5ch-analysis - 5chの過去ログをスクレイピングして、過去流行った単語(ex, 香具師, orz)などを追跡調査
  • tweet_extructor - Twitter日本語評判分析データセットのためのツイートダウンローダ
  • japanese-word-aggregation - Aggregating Japanese words based on Juman++ and ConceptNet5.5
  • jinf - A Japanese inflection converter
  • kwja - A unified language analyzer for Japanese
  • mlm-scoring-transformers - Reproduced package based on Masked Language Model Scoring (ACL2020).
  • ClipCap-for-Japanese - [PyTorch] ClipCap for Japanese
  • SAT-for-Japanese - [PyTorch] Show, Attend and Tell for Japanese
  • cihai - Python library for CJK (Chinese, Japanese, and Korean) language dictionary
  • marine - MARINE : Multi-task leaRnIng-based JapaNese accent Estimation
  • whisper-asr-finetune - Finetuning Whisper ASR model
  • radicalchar - 部首文字正規化ライブラリ
  • akaza - Yet another Japanese IME for IBus/Linux
  • posuto - Japanese postal code data.
  • tacotron2-japanese - Tacotron2 implementation of Japanese
  • ibus-hiragana - ひらがなIME for IBus
  • furiganapad - ふりがなパッド
  • chikkarpy - Japanese synonym library
  • ja-tokenizer-docker-py - Mecab + NEologd + Docker + Python3
  • JapaneseEmbeddingEval - JapaneseEmbeddingEval
  • shuwa - Extend GNOME On-Screen Keyboard for Input Methods
  • japanese-nli-model - This repository provides the code for Japanese NLI model, a fine-tuned masked language model.
  • tra-fugu - A tool for Japanese-English translation and English-Japanese translation by using FuguMT
  • fugumt - ぷるーふおぶこんせぷと で公開した機械翻訳エンジンを利用する翻訳環境です。 フォームに入力された文字列の翻訳、PDFの翻訳が可能です。
  • JaSPICE - JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning Models
  • Retrieval-based-Voice-Conversion-WebUI-JP-localization - jp-localization
  • pyopenjtalk - Python wrapper for OpenJTalk
  • yomigana-ebook - Make learning Japanese easier by adding readings for every kanji in the eBook
  • N46Whisper - Whisper based Japanese subtitle generator
  • japanese_llm_simple_webui - Rinna-3.6B、OpenCALM等の日本語対応LLM(大規模言語モデル)用の簡易Webインタフェースです
  • pdf-translator - pdf-translator translates English PDF files into Japanese, preserving the original layout.
  • japanese_qa_demo_with_haystack_and_es - Haystack + Elasticsearch + wikipedia(ja) を用いた、日本語の質問応答システムのサンプル
  • mozc-devices - Automatically exported from code.google.com/p/mozc-morse
  • vits-japros-webui - 日本語TTS(VITS)の学習と音声合成のGradio WebUI
  • ja-law-parser - A Japanese law parser
  • dictation-kit - Japanese dictation kit using Julius
  • julius4seg - Juliusを使ったセグメンテーション支援ツール
  • voicevox_engine - 無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXの音声合成エンジン
  • LLaVA-JP - LLaVA-JP is a Japanese VLM trained by LLaVA method
  • RAG-Japanese - Open source RAG with Llama Index for Japanese LLM in low resource settting
  • bertjsc - Japanese Spelling Error Corrector using BERT(Masked-Language Model). BERTに基づいて日本語校正
  • llm-leaderboard - Project of llm evaluation to Japanese tasks
  • jglue-evaluation-scripts - Training and evaluation scripts for JGLUE, a Japanese language understanding benchmark
  • BLIP2-Japanese - Modifying LAVIS' BLIP2 Q-former with models pretrained on Japanese datasets.
  • wikipedia-passages-jawiki-embeddings-utils - wikipedia 日本語の文を、各種日本語の embeddings や faiss index へと変換するスクリプト等。
  • simple-simcse-ja - Exploring Japanese SimCSE
  • gpt4-autoeval - GPT-4 を用いて、言語モデルの応答を自動評価するスクリプト
  • t5-japanese - 日本語T5モデル
  • japanese_llm_eval - A repo for evaluating Japanese LLMs ・ 日本語LLMを評価するレポ
  • jmteb - The evaluation scripts of JMTEB (Japanese Massive Text Embedding Benchmark)
  • pydomino - 日本語音声に対して音素ラベルをアラインメントするためのツールです
  • easynovelassistant - 軽量で規制も検閲もない日本語ローカル LLM『LightChatAssistant-TypeB』による、簡単なノベル生成アシスタントです。ローカル特権の永続生成 Generate forever で、当たりガチャを積み上げます。読み上げにも対応。
  • clip-japanese - 日本語データセットでのqlora instruction tuning学習サンプルコード
  • rime-jaroomaji - Japanese rōmaji input schema for Rime IME
  • deep-question-generation - 深層学習を用いたクイズ自動生成(日本語T5モデル)
  • magpie-nemotron - Magpieという手法とNemotron-4-340B-Instructを用いて合成対話データセットを作るコード
  • qlora_ja - 日本語データセットでのqlora instruction tuning学習サンプルコード
  • mozcdic-ut-jawiki - Mozc UT Jawiki Dictionary is a dictionary generated from the Japanese Wikipedia for Mozc.
  • shisa-v2 - Japanese / English Bilingual LLM
  • llm-translator - Mixtral-based Ja-En (En-Ja) Translation model
  • llm-jp-asr - Whisperのデコーダをllm-jp-1.3b-v1.0に置き換えた音声認識モデルを学習させるためのコード
  • rag-japanese - Open source RAG with Llama Index for Japanese LLM in low resource settting
  • monaka - A Japanese Parser (including historical Japanese)
  • jp-translate.cloud - A state-of-the-art open-source Japanese <--> English machine translation system based on the latest NMT research.
  • substring-word-finder - 連続部分文字列の単語判定を行います
  • heron-vlm-leaderboard - This project is a benchmarking tool for evaluating and comparing the performance of various Vision Language Models (VLMs). It uses two datasets: LLaVA-Bench-In-the-Wild and Japanese HERON Bench to measure model performance.
  • text2dataset - Easily turn large English text datasets into Japanese text datasets using open LLMs.
  • mecab-web-api - MeCabを利用した日本語形態素解析WebAPI
  • mecab_controller - Mecab wrapper to generate furigana readings.
  • vits - VITSによるテキスト読み上げ器&ボイスチェンジャー
  • akari_chatgpt_bot - 音声認識、文章生成、音声合成を使って対話するチャットボットアプリ
  • kudasai - Streamlining Japanese-English Translation with Advanced Preprocessing and Integrated Translation Technologies
  • mecab-visualizer - MeCabの形態素解析結果を可視化するツール
  • add-dictionary - OpenJTalkのユーザ辞書をGUIで追加するアプリ
  • j-moshi - J-Moshi: A Japanese Full-duplex Spoken Dialogue System
  • jatts - JATTS: Japanese TTS (for research)
  • tsukasa-speech - a Frontier Japanese Speech Generation net
  • symptom-expression-search - ElasticsearchやGiNZA、患者表現辞書を使った患者表現揺れ吸収する意味構造検索を試した
  • llm-jp-judge - 生成自動評価を行うためのPythonツール
  • asagi-vlm-colaboratory-sample - Colaboratory上でAsagi(合成データセットを活用した大規模日本語VLM)をお試しするサンプル
  • llm-jp-eval-mm - This tool automatically evaluates Japanese multi-modal large language models across multiple datasets.
  • manga109api - Simple python API to read annotation data of Manga109
  • fastrtc-jp - fastrtc用の日本語TTSとSTT追加キット
  • whisper-transcription - Pythonを使用したWhisperモデルによる音声文字起こしツール
  • pocket-researcher - LLMを活用した自律調査エージェント。手軽に情報収集、概要把握。
  • jtransbench - A tool to easily benchmark Japanese translation skills
  • easyllasa - EasyLlasa は 5~15秒の日本語音声と日本語テキストから日本語音声を生成する TSTS (TextSpeechToSpeech) です。
  • kanjikana-model - 氏名漢字カナ突合モデル
  • deep-openreview-research-ja - OpenReview論文を自動で発見・分析する日本語対応AIエージェント
  • pitchbench - Experimental Japanese pitch accent based LLM Benchmark
  • mini-transformer-from-scratch - English to Japanese Transformer from scratch
  • vv_core_inference - VOICEVOXのコア内で用いられているディープラーニングモデルの推論コード
  • pyopenjtalk-plus - pyopenjtalk-plus: A Python wrapper for OpenJTalk with additional improvements
  • japanese_spelling_correction - Japanese Spelling Correction
  • py-kaomoji - python kaomoji
  • llm-jp-vila - This repository contains the code for training llm-jp/llm-jp-3-vila-14b, modified from VILA repository.
  • kanjivg-radical - kanjivg-radical
  • japanese-wordnet-visualization - This project visualizes the Japanese Wordnet (日本語ワードネット) with web application built by Django
  • piper-plus - Enhanced Piper TTS with Japanese support, WebAssembly, multi-GPU training, and quality improvements.
  • Japanera - Easy Tools for Japanese Era System
  • bert-abstractive-text-summarization - Japanese Sentence Summarization with BERT
  • kyujipy - A Python library to convert Japanese texts from Shinjitai (新字体) to Kyujitai (舊字體) and vice versa
  • jitenbot - Web crawler for creating personal copies of Japanese dictionaries
  • ja-icd10 - ICD-10 国際疾病分類の日本語情報を扱うためのPythonパッケージ
  • pl-bert-vits2 - VITS2 using Phoneme-Level Japanese BERT
  • ndc_predictor - NDCPredictorの機械学習モデル(書誌情報から日本十進分類を推測するfastTextの学習済みモデル)
  • pfmt-bench-fin-ja - pfmt-bench-fin-ja: Preferred Multi-turn Benchmark for Finance in Japanese
  • marine-plus - MARINE : Multi-task leaRnIng-based JapaNese accent Estimation (Also supported Windows)
  • ja-tokenizer-benchmark - Compare the speed of various Japanese tokenizers in Python.
  • yat - yat: Yet Another Tokenizer for Japanese NLP
  • igakuqa119 - Evaluating LLMs on the 119th Japanese Medical Licensing Examination
  • japanese-luw-tokenizer - Japanese Long-Unit-Word Tokenizer with RemBertTokenizerFast of Transformers
  • ibus-jig - ibus-jig: Japanese-language Input-method using GPT-4
  • jp-stopword-filter - A lightweight Python library designed to filter stopwords from Japanese text based on customizable rules.
  • yasumail - Synthetic Japanese business email generator for ML training data
  • himotoki - A Python-based Japanese Tokenizer, Dictionary, Morphological Analyzer and Romanization Tool. Based on JMDict for Language Learning.
  • diafill-toolkit - A toolkit for synthesizing filler-rich, short-utterance Japanese dialogue scripts for speech-based interaction using Large Language Models (LLMs) This project is designed to generate data in two phases: Seed Generation (metadata creation) and Dialogue Generation (script creation).
  • eval_vertical_ja - Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
  • jp-llm-corpus-pii-filter - 本コードは,大規模言語モデル(LLM)の学習用コーパスから,個人情報の中でも特に配慮が求められる「要配慮個人情報」をフィルタリングするためのものです.
  • Novel2DialCorpus - 小説テキストから雑談対話コーパスを構築する手法
  • OneCompression - 富士通研究所による LLM 向け後学習量子化 (PTQ) パイプライン。QEP (NeurIPS 2025)、ILP 混合精度、回転前処理、vLLM プラグインを統合。論文: arXiv:2603.28845。
  • manga-translator - Translate text found within speech bubbles in manga images.
  • shirabe-address-api - Shirabe Address API — Japanese address normalization for AI agents (Cloudflare Workers + Fly.io NRT, abr-geocoder backed)
  • medical-paper-summarizer-public - 毎日PubMedから最新の循環器内科論文を自動収集・AI要約してGmailに届けるシステム
  • Irodori-TTS - A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
  • sarashina2.2-tts - Sarashina2.2-TTS is a Japanese-centric text-to-speech system built on a large language model, developed by SB Intuitions.
  • manga-translator - A manga translator created using Yolov8, manga_ocr and deep-translator.
  • jp-tl-bench - Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
  • novel2hermes_jp - メモリ機能が強力なhermes-agentと、日本語検索に強い外部メモリvecmemoriを活かし、長文に耐える小説を企画/プロッティング/執筆するためのskills.md
  • moshi-finetune - Fine-tuning Moshi/J-Moshi on your own spoken dialogue data
  • simple-evals-mm - A lightweight library for evaluating vision language models on English and Japanese benchmarks.
  • medvoice-jp-asr - MedVoice-JP LoRA60: 日本語医療特化ASR (whisper-small + LoRA, 60話者) と国内初の医療ASRベンチマーク (フルFT版 FT60 も収録)
  • open-zeimu-mcp - OSS MCP server for Japanese tax primary sources (法令・通達・タックスアンサー・質疑応答事例・文書回答事例・裁決事例)
  • llm-jp-moshi - LLM-jp-Moshi: Japanese Full-duplex Spoken Dialogue Models
  • shirabe-sdk-python - Official Python SDK for the Shirabe Japan data APIs — ready-made LangChain / OpenAI Agents SDK tools for Japanese name reading (JMnedict-backed candidates), name splitting, address normalization, corporate number lookup, and calendar (rokuyo). Zero-dependency core.
  • fuseji - 日本語特化のPII検出・マスキングミドルウェア(LLMオブザーバビリティ向け)
  • moine - Romanization-aware string comparison for Japanese and Mandarin Chinese.
  • bpe2regex - BPE tokenizer をクソデカ正規表現に変換する意味わからんやつ
  • sanoTTS-jp - 559 K パラメータの日本語 TTS を ESP32-S3 で実時間合成。漢字かな交じり文の形態素解析・アクセント推定まで端末内で走る(M5Stack CoreS3 実機で確認)。推論は依存ゼロの C99、ブラウザ demo あり。arXiv:2608.21378 sanoTTS の日本語 clean-room 再実装。⚠️ コードは MIT ですが、配布モデルの重みは MIT ではありません(LICENSE-MODEL.md。出力に用途制限が伝播します)
  • jev-auto-ime - 打っている言葉が日本語か英語かをJevに聞いて、Macの入力モードを切り替える道具。個人利用のみ。
  • yomiyasu - AI生成の日本語を自然な日本語へ推敲するAgent Skill / Agent Skill for Refining AI-Generated Japanese into Natural Japanese
Name downloads/week total downloads stars last commit
🔗 namedivider-python 📥 1k 📦 110k ⭐ 252 🟡 november 2025
🔗 asa-python 📥 36 📦 32k ⭐ 11 🔴 february 2019
🔗 python_asa - - ⭐ 22 🔴 january 2020
🔗 toiro 📥 51 📦 28k ⭐ 122 🟡 november 2025
🔗 ja-timex 📥 2k 📦 143k ⭐ 140 🔴 november 2023
🔗 JapaneseTokenizers - - ⭐ 138 🔴 invalid
🔗 daaja 📥 40 📦 27k ⭐ 64 🔴 february 2023
🔗 accel-brain-code 📥 228 📦 157k ⭐ 328 🔴 december 2023
🔗 kyoto-reader 📥 1k 📦 69k ⭐ 10 🔴 june 2024
🔗 nlplot 📥 122 📦 117k ⭐ 238 🔴 september 2022
🔗 rake-ja - - ⭐ 21 🔴 october 2018
🔗 jel 📥 25 📦 8k ⭐ 12 🔴 july 2021
🔗 MedNER-J - - ⭐ 18 🔴 may 2022
🔗 zunda-python 📥 10 📦 6k ⭐ 10 🔴 november 2019
🔗 AIO2_DPR_baseline - - ⭐ 16 🔴 january 2022
🔗 showcase 📥 11 📦 8k ⭐ 6 🔴 june 2018
🔗 darts-clone-python 📥 2k 📦 9M ⭐ 20 🔴 april 2022
🔗 jrte-corpus_example - - ⭐ 3 🔴 november 2021
🔗 desuwa 📥 22 📦 11k ⭐ 6 🔴 may 2022
🔗 HotPepperGourmetDialogue - - ⭐ 275 🔴 may 2016
🔗 nlp-recipes-ja - - ⭐ 66 🔴 april 2021
🔗 Japanese_nlp_scripts - - ⭐ 27 🔴 june 2019
🔗 DNorm-J - - ⭐ 9 🔴 june 2022
🔗 pyknp-eventgraph 📥 113 📦 69k ⭐ 9 🔴 september 2022
🔗 ishi 📥 15 📦 7k ⭐ 2 🔴 may 2020
🔗 python-npylm - - ⭐ 34 🔴 january 2019
🔗 python-npycrf - - ⭐ 11 🔴 march 2018
🔗 unsupervised-pos-tagging - - ⭐ 16 🔴 october 2017
🔗 negima 📥 24 📦 17k ⭐ 14 🔴 august 2018
🔗 YouyakuMan - - ⭐ 52 🔴 september 2020
🔗 japanese-numbers-python 📥 1k 📦 2M ⭐ 21 🔴 april 2020
🔗 kantan - - ⭐ 8 🔴 october 2024
🔗 make-meidai-dialogue - - ⭐ 40 🔴 september 2017
🔗 japanese_summarizer - - ⭐ 10 🔴 august 2022
🔗 chirptext 📥 6k 📦 400k ⭐ 7 🔴 october 2022
🔗 yubin 📥 18 📦 3k ⭐ 3 🔴 october 2019
🔗 jawiki-cleaner 📥 43 📦 26k ⭐ 6 🔴 february 2021
🔗 japanese2phoneme 📥 15 📦 5k ⭐ 1 🔴 february 2022
🔗 anlp_nlp2021_d3-1 - - ⭐ 1 🔴 march 2022
🔗 aozora_classification - - ⭐ 11 🔴 september 2017
🔗 aozora-corpus-generator - - ⭐ 9 🔴 june 2025
🔗 JLM - - ⭐ 112 🔴 june 2019
🔗 NTM - - ⭐ 13 🔴 july 2019
🔗 EN-JP-ML-Lexicon - - ⭐ 41 🔴 march 2021
🔗 text-generation - - ⭐ 19 🔴 august 2025
🔗 chainer_nic - - ⭐ 17 🔴 december 2018
🔗 unihan-lm - - ⭐ 2 🔴 november 2020
🔗 mbart-finetuning - - ⭐ 3 🔴 october 2021
🔗 xvector_jtubespeech - - ⭐ 47 🔴 november 2023
🔗 TinySegmenterMaker - - ⭐ 74 🔴 september 2022
🔗 Grongish - - ⭐ 26 🟢 last tuesday
🔗 WordCloud-Japanese - - ⭐ 10 🔴 january 2020
🔗 snark - - ⭐ 11 🔴 march 2020
🔗 toEmoji - - ⭐ 4 🔴 april 2018
🔗 termextract - - ⭐ 18 🔴 september 2018
🔗 JDT-with-KenLM-scoring - - ⭐ 1 🔴 july 2022
🔗 mixture-of-unigram-model - - ⭐ 6 🔴 june 2017
🔗 hidden-markov-model - - ⭐ 5 🔴 june 2017
🔗 Ngram-language-model - - ⭐ 5 🔴 december 2017
🔗 ASRDeepSpeech - - ⭐ 71 🟢 last saturday
🔗 neural_ime - - ⭐ 67 🔴 december 2016
🔗 neural_japanese_transliterator - - ⭐ 178 🔴 september 2017
🔗 tinysegmenter 📥 105k 📦 181k ⭐ repo not found 🔴 november 2015
🔗 AugLy-jp 📥 54 📦 32k ⭐ 7 🟢 september
🔗 furigana4epub 📥 36 📦 13k ⭐ 34 🔴 september 2021
🔗 PyKatsuyou 📥 55 📦 24k ⭐ 13 🔴 march 2025
🔗 jageocoder 📥 5k 📦 482k ⭐ 107 🟡 march
🔗 nksnd - - ⭐ 26 🔴 may 2018
🔗 JaMIE - - ⭐ 9 🟡 march
🔗 fasttext-vs-word2vec-on-twitter-data - - ⭐ 48 🔴 august 2017
🔗 minimal-search-engine - - ⭐ 19 🔴 july 2019
🔗 5ch-analysis - - ⭐ 74 🔴 november 2018
🔗 tweet_extructor - - ⭐ 3 🔴 august 2022
🔗 japanese-word-aggregation - - ⭐ 2 🔴 august 2018
🔗 jinf 📥 590 📦 69k ⭐ 4 🟢 august
🔗 kwja 📥 386 📦 65k ⭐ 147 🟢 september
🔗 mlm-scoring-transformers - - ⭐ 6 🔴 december 2022
🔗 ClipCap-for-Japanese - - ⭐ 13 🔴 october 2022
🔗 SAT-for-Japanese - - ⭐ 2 🔴 october 2022
🔗 cihai 📥 442 📦 229k ⭐ 93 🟢 last saturday
🔗 marine 📥 129 📦 17k ⭐ 38 🔴 september 2022
🔗 whisper-asr-finetune - - ⭐ 32 🔴 december 2022
🔗 radicalchar - - ⭐ 10 🔴 december 2022
🔗 akaza - - ⭐ 264 🟡 june
🔗 posuto 📥 6k 📦 926k ⭐ 236 🟢 last saturday
🔗 tacotron2-japanese - - ⭐ 266 🔴 september 2022
🔗 ibus-hiragana - - ⭐ 81 🟡 march
🔗 furiganapad - - ⭐ 20 🔴 april 2025
🔗 chikkarpy 📥 251 📦 69k ⭐ 55 🔴 february 2022
🔗 ja-tokenizer-docker-py - - ⭐ 36 🔴 may 2022
🔗 JapaneseEmbeddingEval - - ⭐ 184 🔴 october 2024
🔗 shuwa - - ⭐ 145 🔴 december 2022
🔗 japanese-nli-model - - ⭐ 6 🔴 october 2022
🔗 tra-fugu - - ⭐ 5 🔴 march 2023
🔗 fugumt - - ⭐ 65 🔴 february 2021
🔗 JaSPICE 📥 8 📦 2k ⭐ 9 🟡 june
🔗 Retrieval-based-Voice-Conversion-WebUI-JP-localization - - ⭐ 47 🔴 april 2023
🔗 pyopenjtalk 📥 21k 📦 2M ⭐ 261 🔴 april 2025
🔗 yomigana-ebook 📥 29 📦 8k ⭐ 27 🔴 february 2024
🔗 N46Whisper - - ⭐ 1.7k 🔴 february 2025
🔗 japanese_llm_simple_webui - - ⭐ 17 🔴 may 2024
🔗 pdf-translator - - ⭐ 346 🔴 may 2024
🔗 japanese_qa_demo_with_haystack_and_es - - ⭐ 1 🔴 december 2022
🔗 mozc-devices - - ⭐ 2.9k 🟢 last thursday
🔗 vits-japros-webui - - ⭐ 42 🔴 january 2024
🔗 ja-law-parser - - ⭐ 25 🔴 january 2024
🔗 dictation-kit - - ⭐ 166 🔴 april 2019
🔗 julius4seg - - ⭐ 7 🔴 august 2021
🔗 voicevox_engine - - ⭐ 1.8k 🟢 last friday
🔗 LLaVA-JP - - ⭐ 64 🔴 june 2024
🔗 RAG-Japanese - - ⭐ 10 🔴 may 2025
🔗 bertjsc - - ⭐ 14 🔴 august 2024
🔗 llm-leaderboard - - ⭐ 95 🟡 may
🔗 jglue-evaluation-scripts - - ⭐ 18 🟢 last thursday
🔗 BLIP2-Japanese - - ⭐ 14 🔴 september 2025
🔗 wikipedia-passages-jawiki-embeddings-utils - - ⭐ 12 🔴 march 2024
🔗 simple-simcse-ja - - ⭐ 69 🔴 october 2023
🔗 gpt4-autoeval - - ⭐ 17 🔴 june 2024
🔗 t5-japanese - - ⭐ 118 🔴 september 2025
🔗 japanese_llm_eval - - ⭐ 5 🔴 april 2024
🔗 jmteb - - ⭐ 93 🟡 march
🔗 pydomino - - ⭐ 43 🔴 august 2025
🔗 easynovelassistant - - ⭐ 235 🔴 july 2024
🔗 clip-japanese - - ⭐ 13 🔴 september 2025
🔗 rime-jaroomaji - - ⭐ 57 🟢 last thursday
🔗 deep-question-generation - - ⭐ 12 🔴 march 2023
🔗 magpie-nemotron - - ⭐ 8 🔴 july 2024
🔗 qlora_ja - - ⭐ 1 🔴 july 2024
🔗 mozcdic-ut-jawiki - - ⭐ 33 🟢 last friday
🔗 shisa-v2 - - ⭐ 35 🟡 december 2025
🔗 llm-translator - - ⭐ 21 🔴 january 2025
🔗 llm-jp-asr - - ⭐ 9 🔴 september 2024
🔗 rag-japanese - - ⭐ 10 🔴 may 2025
🔗 monaka - - ⭐ 8 🔴 january 2025
🔗 jp-translate.cloud - - ⭐ 3 🔴 september 2024
🔗 substring-word-finder - - ⭐ 4 🟡 november 2025
🔗 heron-vlm-leaderboard - - ⭐ 6 🔴 december 2024
🔗 text2dataset - - ⭐ 32 🔴 january 2025
🔗 mecab-web-api - - ⭐ 40 🔴 july 2022
🔗 mecab_controller - - ⭐ 21 🟢 last friday
🔗 vits - - ⭐ 93 🔴 february 2023
🔗 akari_chatgpt_bot - - ⭐ 48 🟡 october 2025
🔗 kudasai - - ⭐ 26 🔴 june 2025
🔗 mecab-visualizer - - ⭐ 2 🔴 september 2023
🔗 add-dictionary - - ⭐ 3 🟡 october 2025
🔗 j-moshi - - ⭐ 324 🟢 september
🔗 jatts - - ⭐ 44 🟡 march
🔗 tsukasa-speech - - ⭐ 65 🔴 may 2025
🔗 symptom-expression-search - - ⭐ 2 🔴 february 2021
🔗 llm-jp-judge - - ⭐ 63 🟢 september
🔗 asagi-vlm-colaboratory-sample - - ⭐ 1 🔴 march 2025
🔗 llm-jp-eval-mm - - ⭐ 44 🟡 april
🔗 manga109api 📥 74 📦 49k ⭐ 132 🔴 march 2022
🔗 fastrtc-jp - - ⭐ 6 🔴 may 2025
🔗 whisper-transcription - - ⭐ 21 🟡 january
🔗 pocket-researcher - - ⭐ 10 🔴 april 2025
🔗 jtransbench - - ⭐ 13 🟡 october 2025
🔗 easyllasa - - ⭐ 26 🔴 september 2025
🔗 kanjikana-model - - ⭐ 120 🟡 april
🔗 deep-openreview-research-ja - - ⭐ 13 🟡 november 2025
🔗 pitchbench - - ⭐ 2 🟡 february
🔗 mini-transformer-from-scratch - - ⭐ 2 🟡 november 2025
🔗 vv_core_inference - - ⭐ 31 🟡 december 2025
🔗 pyopenjtalk-plus 📥 18k 📦 720k ⭐ 64 🟢 yesterday
🔗 japanese_spelling_correction - - ⭐ 16 🔴 september 2023
🔗 py-kaomoji 📥 42 📦 39k ⭐ 6 🔴 december 2018
🔗 llm-jp-vila - - ⭐ 10 🔴 august 2025
🔗 kanjivg-radical - - ⭐ 107 🔴 august 2018
🔗 japanese-wordnet-visualization - - ⭐ 3 🔴 november 2022
🔗 piper-plus - - ⭐ 218 🟢 last thursday
🔗 Japanera 📥 4k 📦 472k ⭐ 35 🔴 june 2025
🔗 bert-abstractive-text-summarization - - ⭐ 49 🔴 december 2019
🔗 kyujipy 📥 50 📦 24k ⭐ 22 🟡 january
🔗 jitenbot - - ⭐ 4 🔴 december 2024
🔗 ja-icd10 - - ⭐ 5 🔴 july 2021
🔗 pl-bert-vits2 - - ⭐ 14 🔴 december 2023
🔗 ndc_predictor - - ⭐ 14 🔴 august 2021
🔗 pfmt-bench-fin-ja - - ⭐ 9 🔴 march 2025
🔗 marine-plus 📥 142 📦 14k ⭐ 9 🟡 march
🔗 ja-tokenizer-benchmark - - ⭐ 7 🔴 february 2022
🔗 yat - - ⭐ 7 🔴 june 2018
🔗 igakuqa119 - - ⭐ 10 🟢 september
🔗 japanese-luw-tokenizer - - ⭐ 6 🔴 december 2021
🔗 ibus-jig - - ⭐ 4 🔴 december 2023
🔗 jp-stopword-filter 📥 1 📦 7k ⭐ 4 🟢 august
🔗 yasumail - - ⭐ 3 🟡 january
🔗 himotoki 📥 101 📦 7k ⭐ 8 🟢 september
🔗 diafill-toolkit - - ⭐ 0 🟡 january
🔗 eval_vertical_ja - - ⭐ 3 🟡 may
🔗 jp-llm-corpus-pii-filter - - ⭐ 7 🔴 march 2025
🔗 Novel2DialCorpus - - ⭐ 0 🟡 february
🔗 OneCompression - - ⭐ 434 🟢 last thursday
🔗 manga-translator - - ⭐ 27 🟢 august
🔗 shirabe-address-api - - ⭐ 0 🟢 july
🔗 medical-paper-summarizer-public - - ⭐ 20 🟡 april
🔗 Irodori-TTS - - ⭐ 1.4k 🟢 september
🔗 sarashina2.2-tts - - ⭐ 87 🟡 june
🔗 manga-translator - - ⭐ 19 🟡 june
🔗 jp-tl-bench - - ⭐ 6 🟡 february
🔗 novel2hermes_jp - - ⭐ 100 🟢 september
🔗 moshi-finetune - - ⭐ 106 🟡 january
🔗 simple-evals-mm - - ⭐ 5 🟢 august
🔗 medvoice-jp-asr - - ⭐ 5 🟡 june
🔗 open-zeimu-mcp - - ⭐ 1 🟡 june
🔗 llm-jp-moshi - - ⭐ 25 🟡 march
🔗 shirabe-sdk-python - - ⭐ 0 🟢 july
🔗 fuseji - - ⭐ 1 🟡 june
🔗 moine - - ⭐ 10 🟢 september
🔗 bpe2regex - - ⭐ 9 🟢 august
🔗 sanoTTS-jp - - ⭐ 67 🟢 last thursday
🔗 jev-auto-ime - - ⭐ 11 🟢 september
🔗 yomiyasu - - ⭐ 1.5k 🟢 today

C++

Morphology analysis

High-performance libraries for Japanese morphological analysis

  • mecab - Yet another Japanese morphological analyzer
  • jumanpp - Juman++ (a Morphological Analyzer Toolkit)
  • kytea - The Kyoto Text Analysis Toolkit for word segmentation and pronunciation estimation, etc.
  • juman - Japanese Morphological Analysis System JUMAN
Name downloads/week total downloads stars last commit
🔗 mecab - - ⭐ 1.1k 🔴 february 2025
🔗 jumanpp - - ⭐ 413 🟡 april
🔗 kytea - - ⭐ 217 🔴 april 2020
🔗 juman - - ⭐ 12 🔴 december 2021

Parsing

Libraries for dependency and syntactic parsing of Japanese sentences

  • cabocha - Yet Another Japanese Dependency Structure Analyzer
  • knp - A Japanese Parser
Name downloads/week total downloads stars last commit
🔗 cabocha - - ⭐ 120 🔴 february 2025
🔗 knp - - ⭐ 35 🔴 november 2023

Others

Other Japanese NLP and text processing libraries

Show 7 items

  • jsc - Joint source channel model for Japanese Kana Kanji conversion, Chinese pinyin input and CJE mixed input.
  • aquaskk - An input method without morphological analysis.
  • mozc - Mozc - a Japanese Input Method Editor designed for multi-platform
  • trimatch - Trimatch: An (Exact|Prefix|Approximate) String Matching Library
  • resembla - Resembla: Word-based Japanese similar sentence search library
  • corvusskk - ▽▼ SKK-like Japanese Input Method Editor for Windows
  • mozuku - 日本語文章の解析・校正を行う LSP サーバー。
Name downloads/week total downloads stars last commit
🔗 jsc - - ⭐ 15 🔴 december 2012
🔗 aquaskk - - ⭐ 373 🟢 september
🔗 mozc - - ⭐ 3k 🟢 last friday
🔗 trimatch - - ⭐ 2 🟡 february
🔗 resembla - - ⭐ 74 🔴 august 2025
🔗 corvusskk - - ⭐ 376 🟢 august
🔗 mozuku - - ⭐ 419 🟡 april

Rust crate

Morphology analysis

Fast Japanese morphological analysis crates written in Rust

  • lindera - A morphological analysis library.
  • vaporetto - Vaporetto: Very Accelerated POintwise pREdicTion based TOkenizer
  • goya - Japanese Morphological Analysis written in Rust
  • vibrato - vibrato: Viterbi-based accelerated tokenizer
  • yoin - A Japanese Morphological Analyzer written in pure Rust
  • mecab-rs - Safe Rust bindings for mecab a part-of-speech and morphological analyzer library
  • awabi - A morphological analyzer using mecab dictionary
  • kanpyo - Japanese Morphological Analyzer written in Rust
  • mecrab - A pure Rust implementation of a morphological analyzer compatible with MeCab dictionaries (IPADIC format).
Name downloads/week total downloads stars last commit
🔗 lindera - 📦 2.5M ⭐ 685 🟢 yesterday
🔗 vaporetto - 📦 326k ⭐ 299 🟢 july
🔗 goya - 📦 12k ⭐ 84 🔴 december 2021
🔗 vibrato - 📦 95k ⭐ 424 🟢 last saturday
🔗 yoin - 📦 3.1k ⭐ 26 🔴 october 2017
🔗 mecab-rs - 📦 41k ⭐ 72 🔴 september 2023
🔗 awabi - 📦 27k ⭐ 11 🟡 november 2025
🔗 kanpyo - 📦 2.5k ⭐ 109 🟡 february
🔗 mecrab - - ⭐ 8 🟡 january

Converter

Crates for script and character conversion in Japanese text

  • wana_kana_rust - Utility library for checking and converting between Japanese characters - Hiragana, Katakana - and Romaji
  • unicode-jp-rs - A Rust library to convert Japanese Half-width-kana[半角カナ] and Wide-alphanumeric[全角英数] into normal ones
  • kana - [Mirror] CLI program for transliterating romaji text to either hiragana or katakana
  • kanaria - このライブラリは、ひらがな・カタカナ、半角・全角の相互変換や判別を始めとした機能を提供します。
  • japanese-address-parser - 日本の住所を都道府県/市区町村/町名/その他に分割するライブラリです
  • yosina - Yosina is a transliteration library deals with the letters and symbols used in Japanese writing.
  • mojimoji-rs - Rust implementation of a fast converter between Japanese hankaku and zenkaku characters, mojimoji.
  • haqumei - A Japanese Grapheme-to-Phoneme (G2P) library.
  • ja-furigana - 日本語フリガナ (ルビ) を扱う Rust 製ライブラリ + ローカル HTTP サーバー。ルールはすべてデータ駆動 (TOML)。
  • jpnorm - 日本語テキスト正規化ライブラリ (Rust core + Python)。neologdn 互換・用途別プリセット・URL 保護・カスタム辞書 / Fast, configurable Japanese text normalization
Name downloads/week total downloads stars last commit
🔗 wana_kana_rust - 📦 633k ⭐ 91 🟡 may
🔗 unicode-jp-rs - 📦 70k ⭐ 19 🔴 april 2020
🔗 kana - - ⭐ 12 🔴 january 2023
🔗 kanaria - - ⭐ 21 🟡 february
🔗 japanese-address-parser - - ⭐ 14 🟡 june
🔗 yosina - - ⭐ 27 🟡 april
🔗 mojimoji-rs - - ⭐ 4 🔴 november 2022
🔗 haqumei - - ⭐ 10 🟢 september
🔗 ja-furigana - - ⭐ 21 🟢 last friday
🔗 jpnorm - - ⭐ 1 🟢 september

Search engine library

Libraries for Japanese full-text search and indexing

  • lindera-tantivy - Lindera tokenizer for Tantivy.
  • tantivy-vibrato - A Tantivy tokenizer using Vibrato.
  • sqlite-vaporetto - SQLite FTS5 extension for fast Japanese full-text search with 🛥Vaporetto / Vaporetto による高速な日本語全文検索を SQLite FTS5 で実現する拡張機能
  • duckdb-vaporetto - DuckDB extension for Japanese full-text search with 🛥Vaporetto / Vaporetto による DuckDB + 日本語全文検索拡張機能
  • lindera-sqlite - Lindera for SQLite FTS5 extention
Name downloads/week total downloads stars last commit
🔗 lindera-tantivy - 📦 230k ⭐ 70 🟢 august
🔗 tantivy-vibrato - 📦 1.6k ⭐ 3 🟡 april
🔗 sqlite-vaporetto - - ⭐ 21 🟡 april
🔗 duckdb-vaporetto - - ⭐ 12 🟡 april
🔗 lindera-sqlite - - ⭐ 23 🟢 august

Others

Supplementary crates for Japanese text and IME processing

Show 26 items

  • daachorse - A fast implementation of the Aho-Corasick algorithm using the compact double-array data structure in Rust.
  • find-simdoc - Finding all pairs of similar documents time- and memory-efficiently
  • crawdad - Rust library of natural language dictionaries using character-wise double-array tries.
  • tokenizer-speed-bench - Comparison code of various tokenizers
  • stringmatch-bench - Here provides benchmark tools to compare the performance of data structures for string matching.
  • vime - Using Vim as an input method for X11 apps
  • voicevox_core - 無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのコア
  • akaza - Yet another Japanese IME for IBus/Linux
  • Jotoba - A free online, self-hostable, multilang Japanese dictionary.
  • dvorakjp-romantable - Google 日本語入力用DvorakJPローマ字テーブル / DvorakJP Roman Table for Google Japanese Input
  • niinii - Japanese glossator for assisted reading of text using Ichiran
  • cskk - SKK (Simple Kana Kanji henkan) library
  • japanki - Learn Japanese vocabs 🇯🇵 by doing quizzes on CLI!
  • jpreprocess - Japanese text preprocessor for Text-to-Speech applications (OpenJTalk rewrite in rust language)
  • listup_precedent - 裁判例のデータ一覧を裁判所のホームページ(https://www.courts.go.jp/index.html) をスクレイピングして生成するソフトウェア
  • jisho - Jisho is a CLI tool & Rust library that provides a Japanese-English dictionary.
  • kanalizer - 英単語から読みを推測するライブラリ。
  • koharu - Automated manga translation tool with LLM, written in Rust.
  • yomine - A Japanese vocabulary mining tool designed to help language learners mine new words and expressions.
  • matsuba - lightweight japanese ime written in rust
  • hujiang_dictionary - 日本語辞書 by Rust, support Telegram bot, AWS Lambda and Cloudflare Workers. Support LLM and search RAG.
  • mecab-dic-converter - MeCab 用のコンパイル済み辞書を読み取り、解析・再構成し、最終的に Vibrato や Lindera のようなMecab互換 tokenizer が使える辞書へ変換するための Rust crate です。
  • jp-deinflector - A high-performance Rust crate for deinflecting Japanese words using perfect hash tables
  • mikke - 日本語 Markdown ノートのローカル検索 CLI 👀 — BM25 全文検索 (SQLite FTS5) + optional なローカル semantic/hybrid。単一バイナリ・外部 API 不使用、AI コーディングエージェント向け。
  • suiko - 日本語文書の自然さと読みやすさを再現可能に診断するRust CLI / Deterministic diagnostics for natural and readable Japanese writing
  • daac-bpe - Fast incremental BPE tokenizer based on daachorse
Name downloads/week total downloads stars last commit
🔗 daachorse - 📦 7M ⭐ 284 🟢 august
🔗 find-simdoc - 📦 29k ⭐ 62 🔴 march 2025
🔗 crawdad - 📦 160k ⭐ 39 🟢 august
🔗 tokenizer-speed-bench - - ⭐ 4 🔴 march 2023
🔗 stringmatch-bench - - ⭐ 3 🔴 september 2022
🔗 vime - - ⭐ 228 🔴 november 2022
🔗 voicevox_core - - ⭐ 1.1k 🟢 yesterday
🔗 akaza - - ⭐ 264 🟡 june
🔗 Jotoba - - ⭐ 210 🔴 january 2024
🔗 dvorakjp-romantable - - ⭐ 58 🟢 last thursday
🔗 niinii - - ⭐ 16 🟢 last saturday
🔗 cskk - - ⭐ 85 🟢 september
🔗 japanki - - ⭐ 4 🔴 october 2023
🔗 jpreprocess - - ⭐ 60 🟡 june
🔗 listup_precedent - - ⭐ 7 🟢 july
🔗 jisho - - ⭐ 18 🟡 april
🔗 kanalizer - - ⭐ 31 🟡 may
🔗 koharu - - ⭐ 5.7k 🟢 today
🔗 yomine - - ⭐ 65 🟢 last saturday
🔗 matsuba - - ⭐ 19 🔴 march 2023
🔗 hujiang_dictionary - - ⭐ 72 🟢 today
🔗 mecab-dic-converter - - ⭐ 1 🟡 may
🔗 jp-deinflector - - ⭐ 7 🟢 july
🔗 mikke - - ⭐ 49 🟢 august
🔗 suiko - - ⭐ 115 🟢 last tuesday
🔗 daac-bpe - - ⭐ 3 🟢 august

JavaScript

Morphology analysis

Japanese morphological analysis libraries for browser and Node.js

  • kuromoji.js - JavaScript implementation of Japanese morphological analyzer
  • rakutenma - Rakuten MA - morphological analyzer (word segmentor + PoS Tagger) for Chinese and Japanese written purely in JavaScript.
  • node-mecab-ya - Yet another mecab wrapper for nodejs
  • juman-bin - a User-Extensible Morphological Analyzer for Japanese. 日本語形態素解析システム
  • node-mecab-async - Asynchronous japanese morphological analyser using MeCab.
Name downloads/week total downloads stars last commit
🔗 kuromoji.js 📥 373k/week 📦 13M ⭐ 1k 🔴 november 2018
🔗 rakutenma 📥 19/week 📦 1.3k ⭐ 471 🔴 january 2015
🔗 node-mecab-ya 📥 306/week 📦 8.1k ⭐ 110 🔴 june 2021
🔗 juman-bin 📥 4/week 📦 487 ⭐ 3 🔴 may 2017
🔗 node-mecab-async 📥 2.9k/week 📦 318k ⭐ 104 🔴 october 2017

Converter

Libraries for converting Japanese scripts and readings

  • kuroshiro - Japanese language library for converting Japanese sentence to Hiragana, Katakana or Romaji with furigana and okurigana modes supported.
  • kuroshiro-analyzer-kuromoji - Kuromoji morphological analyzer for kuroshiro.
  • hepburn - Node.js module for converting Japanese Hiragana and Katakana script to, and from, Romaji using Hepburn romanisation
  • japanese-numerals-to-number - Converts Japanese Numerals into number
  • jslingua - Javascript libraries to process text: Arabic, Japanese, etc.
  • WanaKana - Javascript library for detecting and transliterating Hiragana <--> Katakana <--> Romaji
  • node-romaji-name - Normalize and fix common issues with Romaji-based Japanese names.
  • kyujitai.js - Utility collections for making Japanese text old-fashioned
  • normalize-japanese-addresses - オープンソースの住所正規化ライブラリ。
  • jaconv - 日本語文字変換ライブラリ (javascript)
  • romaji-conv - Convert romaji into hiragana
  • japanese-addresses-v2 - 全国の住所データAPI
  • jptext-to-emoji - テキストの単語を絵文字に変換する
  • japanese.js - Util collection for Japanese text processing. Hiraganize, Katakanize, and Romanize.
  • genshijin - About genshijin 原始人 🗿| Claude Code / Codex等AIエージェント 向け超圧縮コミュニケーションスキル。caveman の日本語版をベースに、日本語特有の冗長表現に最適化。
Name downloads/week total downloads stars last commit
🔗 kuroshiro 📥 27k/week 📦 911k ⭐ 995 🟢 today
🔗 kuroshiro-analyzer-kuromoji 📥 27k/week 📦 891k ⭐ 72 🟢 september
🔗 hepburn 📥 75k/week 📦 6.7M ⭐ 131 🟢 july
🔗 japanese-numerals-to-number 📥 118k/week 📦 3.6M ⭐ 59 🔴 february 2023
🔗 jslingua 📥 99/week 📦 9.7k ⭐ 53 🔴 october 2023
🔗 WanaKana 📥 106k/week 📦 3.6M ⭐ 943 🟡 june
🔗 node-romaji-name 📥 142/week 📦 22k ⭐ 39 🔴 december 2023
🔗 kyujitai.js 📥 123/week 📦 3.3k ⭐ 24 🔴 august 2020
🔗 normalize-japanese-addresses - - ⭐ 963 🟢 september
🔗 jaconv - - ⭐ 87 🔴 june 2025
🔗 romaji-conv - - ⭐ 25 🟡 february
🔗 japanese-addresses-v2 - - ⭐ 84 🟢 september
🔗 jptext-to-emoji - - ⭐ 1 🟡 june
🔗 japanese.js - - ⭐ 167 🔴 august 2020
🔗 genshijin - - ⭐ 332 🟢 august

Others

Other libraries for Japanese NLP in JavaScript

Show 26 items

  • bangumi-data - Raw data for Japanese Anime
  • yomichan - Japanese pop-up dictionary extension for Chrome and Firefox.
  • proofreading-tool - GUIで動作する文書校正ツール GUI tool for textlinting.
  • kanjigrid - A web-app displaying the 2200 kanji characters taught in James Heisig's "Remembering the Kanji", 6th edition.
  • japanese-toolkit - Monorepo for Kanji, Furigana, Japanese DB, and others
  • analyze-desumasu-dearu - 文の敬体(ですます調)、常体(である調)を解析するJavaScriptライブラリ
  • hatsuon - Japanese pitch accent utils
  • sentiment_ja_js - Sentiment Analysis in Japanese. sentiment_ja with JavaScript
  • mecab-ipadic-seed - mecab-ipadic seed dictionary reader
  • oskim - Extend GNOME On-Screen Keyboard for Input Methods
  • tweetMapping - 東日本大震災発生から24時間以内につぶやかれたジオタグ付きツイートのデジタルアーカイブです。
  • pitch-accent - Predict pitch accent in Japanese
  • kana2ipa - 「ひらがな」または「カタカナ」を日本語で発音する際の音声記号(IPA)に変換するコマンド
  • voicevox - 無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのエディター
  • kamiya-codec - Towards a Japanese verb conjugator and deconjugator based on Taeko Kamiya's The Handbook of Japanese Verbs and The Handbook of Japanese Adjectives and Adverbs opuses.
  • closewords - 最も似た単語を単語群から検索する日本語(漢字含む)対応のライブラリ
  • japanese-analyzer - Japanese Sentence Analyzer (日本語文章解析器)
  • japanese-furigana-normalize - Normalize Japanese Furigana
  • yama - acquire Japanese vocabulary on any website
  • kaitai - An application for analyzing Japanese sentence structure using AI. This tool visualizes how words and phrases relate to each other, showing grammatical relationships with interactive diagrams.
  • tsukeru-furigana-converter - Browser extension (Chrome/Edge/Firefox) that injects furigana into Japanese webpages on-demand; includes dictionary tooltips, JLPT filtering, and vocab/Anki export.
  • sudachi-synonyms-dictionary - Sudachi's synonyms dictionary
  • qmd-ja - Japanese-enhanced fork of qmd — Vaporetto WASM morphological tokenizer for accurate Japanese BM25 search
  • shirabe-sdk - Official TypeScript SDK for the Shirabe Japan data APIs — ready-made Vercel AI SDK / LangChain tools for Japanese name splitting/reading, address normalization, corporate number lookup, and calendar (rokuyo). Zero runtime dependencies in the core.
  • pii-ja-ner-onnx-demo - PII-JA NER Browser Demo
  • jev-semgrep - grep by meaning, across languages. TypeSafe Jev scores every line against a meaning; combine meanings with AND/OR/NOT. 意味で探す grep。日本語で英語を、英語で日本語を検索できる
Name downloads/week total downloads stars last commit
🔗 bangumi-data 📥 1k/week 📦 63k ⭐ 648 🟢 yesterday
🔗 yomichan - - ⭐ 1.1k 🔴 february 2023
🔗 proofreading-tool - - ⭐ 87 🟡 october 2025
🔗 kanjigrid - - ⭐ 45 🔴 november 2018
🔗 japanese-toolkit - - ⭐ 64 🔴 january 2023
🔗 analyze-desumasu-dearu 📥 287k/week 📦 7.9M ⭐ 18 🟡 april
🔗 hatsuon 📥 44/week 📦 2.3k ⭐ 38 🔴 march 2022
🔗 sentiment_ja_js - - ⭐ 10 🔴 december 2021
🔗 mecab-ipadic-seed 📥 373/week 📦 9.9k ⭐ 8 🔴 july 2016
🔗 oskim - - ⭐ 2 🔴 february 2023
🔗 tweetMapping - - ⭐ 27 🟡 march
🔗 pitch-accent 📥 3/week 📦 198 ⭐ 2 🔴 september 2023
🔗 kana2ipa - - ⭐ 19 🔴 october 2020
🔗 voicevox - - ⭐ 3.3k 🟢 last saturday
🔗 kamiya-codec - - ⭐ 29 🔴 may 2025
🔗 closewords - - ⭐ 5 🟡 may
🔗 japanese-analyzer - - ⭐ 802 🟢 yesterday
🔗 japanese-furigana-normalize - - ⭐ 8 🔴 july 2024
🔗 yama - - ⭐ 8 🟡 february
🔗 kaitai - - ⭐ 1 🟢 september
🔗 tsukeru-furigana-converter - - ⭐ 9 🟢 september
🔗 sudachi-synonyms-dictionary - - ⭐ 15 🟢 july
🔗 qmd-ja - - ⭐ 2 🟢 september
🔗 shirabe-sdk - - ⭐ 0 🟢 july
🔗 pii-ja-ner-onnx-demo - - ⭐ 0 🟢 july
🔗 jev-semgrep - - ⭐ 146 🟢 today

Go

Morphology analysis

Lightweight Japanese morphological analysis libraries in Go

  • kagome - Self-contained Japanese Morphological Analyzer written in pure Go
Name downloads/week total downloads stars last commit
🔗 kagome - - ⭐ 984 🟢 september

Others

Additional Go-based Japanese text processing libraries

Show 9 items

  • ojosama - テキストを壱百満天原サロメお嬢様風の口調に変換します
  • nihongo - Japanese Dictionary
  • yomichan-import - External dictionary importer for Yomichan.
  • imas-ime-dic - THE IDOLM@STER words dictionary for Japanese IME (by imas-db.jp)
  • go-kakasi - Kanji transliteration to hiragana/katakana/romaji, in Go
  • go-moji - A Go library for Zenkaku/Hankaku conversion
  • ojichat - おじさんがLINEやメールで送ってきそうな文を生成する
  • name - Name Searcher in Japanese
  • jp-pii-detector - 日本語個人情報検出器
Name downloads/week total downloads stars last commit
🔗 ojosama - - ⭐ 387 🟢 september
🔗 nihongo - - ⭐ 83 🟡 june
🔗 yomichan-import - - ⭐ 88 🔴 february 2023
🔗 imas-ime-dic - - ⭐ 32 🟡 january
🔗 go-kakasi - - ⭐ 6 🟢 september
🔗 go-moji - - ⭐ 21 🔴 april 2019
🔗 ojichat - - ⭐ 1.3k 🔴 october 2024
🔗 name - - ⭐ 11 🔴 january 2025
🔗 jp-pii-detector - - ⭐ 4 🟢 august

Java

Morphology analysis

Japanese morphological analysis and dictionary management libraries

  • kuromoji - Kuromoji is a self-contained and very easy to use Japanese morphological analyzer designed for search
  • Sudachi - A Japanese Tokenizer for Business
  • SudachiDict - A lexicon for Sudachi
  • meval - 形態素解析器性能評価システム MevAL
Name downloads/week total downloads stars last commit
🔗 kuromoji - - ⭐ 1.1k 🔴 september 2019
🔗 Sudachi - - ⭐ 1k 🟢 september
🔗 SudachiDict - - ⭐ 311 🟢 september
🔗 meval - - ⭐ 7 🔴 august 2019

Others

Java libraries for Japanese NLP and OCR

Show 9 items

  • kanjitomo-ocr - Java library for identifying Japanese characters from images
  • jakaroma - Java library and command-line tool to transliterate Japanese kanji to romaji (Latin alphabet)
  • kakasi-java - Kanji transliteration to hiragana/katakana/romaji, in Java
  • Kamite - A desktop language immersion companion for learners of Japanese
  • react-native-japanese-tokenizer - Async Japanese Tokenizer Native Plugin for React Native for iOS and Android
  • elasticsearch-analysis-japanese - Japanese analyzer uses kuromoji japanese tokenizer for ElasticSearch
  • moji4j - A Java library to converts between Japanese Hiragana, Katakana, and Romaji scripts.
  • neologdn-java - Japanese text normalizer for mecab-neologd
  • elasticsearch-sudachi - The Japanese analysis plugin for elasticsearch
Name downloads/week total downloads stars last commit
🔗 kanjitomo-ocr - - ⭐ 206 🔴 may 2021
🔗 jakaroma - - ⭐ 71 🔴 june 2025
🔗 kakasi-java - - ⭐ 56 🔴 april 2016
🔗 Kamite - - ⭐ 144 🔴 march 2025