Natuurlijke taalverwerking
77 projecten in AI & ML
sentence-transformers
@UKPLabState-of-the-Art Embeddings, Retrieval, and Reranking
unstructured
@Unstructured-IOConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
LanguageTool
@languagetool-orgProofread more than 20 languages. It finds many errors that a simple spell checker cannot detect.
BettaFish
@666ghj微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
bible
@thiagobodrukBible text in JSON and XML — 90 versions across 35 languages, ready for apps, verse search, and NLP/AI datasets
Ciphey
@bee-san⚡ Automatically decrypt encryptions without knowing the key or cipher, decode encodings, and crack hashes ⚡
nltk
@nltkNLTK Source
tokenizers
@huggingface💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
presidio
@data-privacy-stackAn open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data. Supports NLP, pattern matching, and customizable pipelines.
sie
@superlinkedOpen-source inference server and production cluster for all the models your agent needs.
MOSS
@OpenMOSSAn open-source, tool-augmented conversational language model from Fudan University
compromise
@spencermountainmodest natural-language processing
sentencepiece
@googleUnsupervised text tokenizer for Neural Network-based text generation.
gensim
@piskvorkyTopic Modelling for Humans
flash-linear-attention
@fla-org🚀 Efficient implementations for emerging model architectures
awesome-legal-data
@openlegaldataA collection of datasets and other resources for legal text processing.
lingvo
@tensorflowLingvo
MNBVC
@esbatmopMNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。
RAGHub
@Andrew-JangA community-driven collection of RAG (Retrieval-Augmented Generation) frameworks, projects, and resources. Contribute and explore the evolving RAG ecosystem.
CoreNLP
@stanfordnlpCoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
TextBlob
@sloriaSimple, Pythonic, text processing--Sentiment analysis, part-of-speech tagging, noun phrase extraction, translation, and more.
argilla
@argilla-ioArgilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
HanLP
@hankcs中文分词 词性标注 命名实体识别 依存句法分析 成分句法分析 语义依存分析 语义角色标注 指代消解 风格转换 语义相似度 新词发现 关键词短语提取 自动摘要 文本分类聚类 拼音简繁转换 自然语言处理
rasa
@RasaHQ💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
awesome-bioie
@caufieldjh🧫 A curated list of resources relevant to doing Biomedical Information Extraction (including BioNLP)
scispacy
@allenaiA full spaCy pipeline and models for scientific/biomedical documents.
sd
@chmlnIntuitive find & replace CLI (sed alternative)
awesome-search
@frutikAwesome Search - this is all about the (e-commerce, but not only) search and its awesomeness
shekar
@amirivojdanSimplifying Persian NLP for Modern Applications
TagUI
@aisingaporeFree RPA tool by AI Singapore
argos-translate
@argosopentechOpen-source offline translation library written in Python
nlp_chinese_corpus
@brightmart大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
ctakes
@apacheApache cTAKES is a Natural Language Processing (NLP) platform for clinical text.
storm
@stanford-ovalAn LLM-powered knowledge curation system that researches a topic and generates a full-length report with citations.
autoprompt
@ucinlpAutoPrompt: Automatic Prompt Construction for Masked Language Models.
Deta_Parser
@yaoguangluo快速中文分词分析word segmentation
LLMsPracticalGuide
@Mooler0410A curated list of practical guide resources of LLMs (LLMs Tree, Examples, Papers)
sensitive-word
@houbb👮♂️The sensitive word tool for java.(敏感词/违禁词/违法词/脏词。基于 DFA 算法实现的高性能 java 敏感词过滤工具框架。内置支持单词标签分类分级。请勿发布涉及政治、广告、营销、翻墙、违反国家法律法规等内容。高性能敏感词检测过滤组件,附带繁体简体互换,支持全角半角互换,汉字转拼音,模糊搜索等功能。)
obsei
@obseiObsei is a low code AI powered automation tool. It can be used in various business flows like social listening, AI based alerting, brand image analysis, comparative study and more .
fuzi.mingcha
@irlab-sdu夫子•明察司法大模型是由山东大学、浪潮云、中国政法大学联合研发,以 ChatGLM 为大模型底座,基于海量中文无监督司法语料与有监督司法微调数据训练的中文司法大模型。该模型支持法条检索、案例分析、三段论推理判决以及司法对话等功能,旨在为用户提供全方位、高精准的法律咨询与解答服务。
gr-nlp-toolkit
@nlpauebThe Greek NLP toolkit for Python. Supports NER/DP/POS Tagging/Greeklish-to-Greek Transliteration. Visit the web demo here: https://huggingface.co/spaces/AUEB-NLP/greek-nlp-toolkit-demo (paper presented at COLING 2025)
NLP-progress
@sebastianruderRepository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
Unix-Text-Processing
@larrykollarRecreated sources for the book "UNIX Text Processing," published in 1987.
models
@PaddlePaddleOfficially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.
nlp.js
@axa-groupAn NLP library for building bots, with entity extraction, sentiment analysis, automatic language identify, and so more
flashtext
@vi3k6i5Extract Keywords from sentence or Replace keywords in sentences.
flair
@flairNLPA very simple framework for state-of-the-art Natural Language Processing (NLP)
extractous
@yobix-aiFast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
fashion-clip
@patrickjohncyhFashionCLIP is a CLIP-like model fine-tuned for the fashion domain.
Legal-Text-Analytics
@Liquid-Legal-InstituteA list of selected resources, methods, and tools dedicated to Legal Text Analytics.
ltp
@HIT-SCIRLanguage Technology Platform
Med-ChatGLM
@SCIR-HIRepo for Chinese Medical ChatGLM 基于中文医学知识的ChatGLM指令微调
browser-ml-inference
@jobergumEdge Inference in Browser with Transformer NLP model
Awesome-Chinese-NLP
@crownpkuA curated list of resources for Chinese NLP 中文自然语言处理相关资料
GPT2-Chinese
@MorizeyaoChinese version of GPT2 training code, using BERT tokenizer.
BERT-pytorch
@codertimoGoogle AI 2018 BERT pytorch implementation
Baichuan-7B
@baichuan-incA large-scale 7B pretraining language model developed by BaiChuan-Inc.
mordecai
@openeventdataFull text geoparsing as a Python library
ansj_seg
@NLPchinaansj分词.ict的真正java实现.分词效果速度都超过开源版的ict. 中文分词,人名识别,词性标注,用户自定义词典
Leaf-Question-Generation
@KristiyanVachevEasy to use and understand multiple-choice question generation algorithm using T5 Transformers.
ailearning
@apachecnAiLearning:数据分析+机器学习实战+线性代数+PyTorch+NLTK+TF2
Jiagu
@ownthinkJiagu深度学习自然语言处理工具 知识图谱关系抽取 中文分词 词性标注 命名实体识别 情感分析 新词发现 关键词 文本摘要 文本聚类
awesome_Chinese_medical_NLP
@GanjinZero中文医学NLP公开资源整理:术语集/语料库/词向量/预训练模型/知识图谱/命名实体识别/QA/信息抽取/模型/论文/etc
Sharetape-Open-Source
@adhikary97Script that takes any long form video or podcast and outputs clips for social media
AILA-Artificial-Intelligence-for-Legal-Assistance
@Ananyapam7Python implementations of the various methods used in FIRE 2019 conference.
pytorch-pos-tagging
@bentrevettA tutorial on how to implement models for part-of-speech tagging using PyTorch and TorchText.
stanford-nlp-tagger
@patrickschurPHP wrapper for the Stanford Natural Language Processing library. Supports POSTagger and CRFClassifier.
wink-pos-tagger
@winkjsEnglish Part-of-speech (POS) tagger
ParseLawDocuments
@yinhao0214对收集的法律文档进行一系列分析,包括根据规范自动切分、案件相似度计算、案件聚类、法律条文推荐等(试验目前基于婚姻类案件,可扩展至其它领域)。
WantWords
@thunlpAn open-source online reverse dictionary.
monpa
@monpa-teamMONPA 罔拍是一個提供正體中文斷詞、詞性標註以及命名實體辨識的多任務模型
ner-slot_filling
@GaoQ1中文自然语言的实体抽取和意图识别(Natural Language Understanding),可选Bi-LSTM + CRF 或者 IDCNN + CRF
Natural-Language-Processing-Specialization
@amanjeetsahuThis repo contains my coursework, assignments, and Slides for Natural Language Processing Specialization by deeplearning.ai on Coursera
knowledge_graph
@AnjaneyaTripathiKnowledge Graph for Legal Documents using Litigation Releases from the SEC website. Classifies into different crimes, extracts relevant information (violator, violation, action taken by authorities and fines) and finally generates the knowledge graph.
cogcomp-nlp
@CogCompCogComp's Natural Language Processing Libraries and Demos: Modules include lemmatizer, ner, pos, prep-srl, quantifier, question type, relation-extraction, similarity, temporal normalizer, tokenizer, transliteration, verb-sense, and more.
BillSum
@FiscalNoteUS Bill Summarization Corpus
Fake-Reviews-Detection
@SayamAltSuccessfully developed a machine learning model which can predict whether an online review is fraudulent or not. The main idea used to detect the fake nature of reviews is that the review should be computer generated through unfair means. If the review is created manually, then it is considered legal and original.