Keyword Extraction Explorer Tool
Extract the most important keywords and phrases from your text using various algorithms! This tool compares statistical methods (YAKE, TF-IDF) with graph-based methods (TopicRank, MultipartiteRank, PositionRank, TextRank) so you can see how each approaches the same text.
How to use:
- ๐ Enter your text in the text area below
- ๐ฏ Select a model from the dropdown for keyword extraction
- โ๏ธ Adjust parameters (number of keywords, n-gram range)
- ๐ Click "Extract Keywords" to see results with organized output
โน๏ธ Model Descriptions
All six methods here are statistical or graph-based. They need no training data and work directly from the words in your text. They differ in how they decide which words matter.
- YAKE (statistical):
- Scores each word using simple text features — how often it appears, where it sits, whether it's capitalised, and what surrounds it. Good all-rounder, works on short texts and many languages.
- TF-IDF (statistical):
- Picks out words that are common in your text but rare in general — the ones that make it distinctive. A long-standing, dependable baseline.
- TopicRank (graph-based):
- Groups similar words into topics first, then ranks the topics and picks one phrase to represent each. Reduces near-duplicates, so results cover a broader spread of themes.
- MultipartiteRank (graph-based):
- A refinement of TopicRank that keeps individual phrases as the unit while still respecting topics. Tends to give the most balanced, precise keyphrases (the strongest performer in our tests).
- PositionRank (graph-based):
- Gives extra weight to words that appear early in the text, on the idea that important terms are often introduced up front. Useful for structured writing like articles and abstracts.
- TextRank (graph-based):
- Builds a network of words that appear near each other and ranks them using the same idea as Google's PageRank — a word is important if other important words connect to it.
TopicRank, MultipartiteRank, PositionRank and TextRank are graph-based methods provided by the pke library.
๐ Detailed Results
| ๐ Text to Analyse | ๐ฏ Select Keyword Extraction Model | ๐ Number of Keywords | Min N-gram | Max N-gram |
|---|
๐ Model Information & Documentation
Learn more about the algorithms used in this tool:
- YAKE: Yet Another Keyword Extractor โ
- TF-IDF: Term Frequency-Inverse Document Frequency โ
- TopicRank: Bougouin et al. (2013) โ
- MultipartiteRank: Boudin (2018) โ
- PositionRank: Florescu & Caragea (2017) โ
- TextRank: Mihalcea & Tarau (2004) โ
- pke library (TopicRank, MultipartiteRank, PositionRank, TextRank): Python Keyphrase Extraction โ
This Keyword Extraction Explorer Tool was created as part of the Digital Scholarship at Oxford (DiSc) funded research project: Extracting Keywords from Crowdsourced Collections.
See also the main EKCC project repository: Crowdsourced Data Tools โ
and Arana-Catania, M., Conisbee, C., & Kidd, M. (2026). Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI. https://doi.org/10.48550/arXiv.2607.09324 โ
The code for this tool was built with the aid of Claude Opus 4 and updated with Opus 4.8.