Keyword Extraction Explorer Tool

Extract the most important keywords and phrases from your text using various algorithms! This tool compares statistical methods (YAKE, TF-IDF) with graph-based methods (TopicRank, MultipartiteRank, PositionRank, TextRank) so you can see how each approaches the same text.

How to use:

  1. ๐Ÿ“ Enter your text in the text area below
  2. ๐ŸŽฏ Select a model from the dropdown for keyword extraction
  3. โš™๏ธ Adjust parameters (number of keywords, n-gram range)
  4. ๐Ÿ” Click "Extract Keywords" to see results with organized output
๐Ÿ’ก Top tip: Different models excel at different types of texts - experiment to find the best one for your content!
๐ŸŽฏ Select Keyword Extraction Model
5 30
1 3
1 4
๐Ÿ’ก Top tip: N-grams are sequences of words. Set Min=1, Max=3 to extract single words, phrases of 2 words, and phrases of 3 words. Higher values capture longer phrases but may reduce precision.
โ„น๏ธ Model Descriptions

All six methods here are statistical or graph-based. They need no training data and work directly from the words in your text. They differ in how they decide which words matter.

YAKE
(statistical):
Scores each word using simple text features — how often it appears, where it sits, whether it's capitalised, and what surrounds it. Good all-rounder, works on short texts and many languages.
TF-IDF
(statistical):
Picks out words that are common in your text but rare in general — the ones that make it distinctive. A long-standing, dependable baseline.
TopicRank
(graph-based):
Groups similar words into topics first, then ranks the topics and picks one phrase to represent each. Reduces near-duplicates, so results cover a broader spread of themes.
MultipartiteRank
(graph-based):
A refinement of TopicRank that keeps individual phrases as the unit while still respecting topics. Tends to give the most balanced, precise keyphrases (the strongest performer in our tests).
PositionRank
(graph-based):
Gives extra weight to words that appear early in the text, on the idea that important terms are often introduced up front. Useful for structured writing like articles and abstracts.
TextRank
(graph-based):
Builds a network of words that appear near each other and ranks them using the same idea as Google's PageRank — a word is important if other important words connect to it.

TopicRank, MultipartiteRank, PositionRank and TextRank are graph-based methods provided by the pke library.

๐Ÿ“‹ Detailed Results

Examples
๐Ÿ“ Text to Analyse ๐ŸŽฏ Select Keyword Extraction Model ๐Ÿ“Š Number of Keywords Min N-gram Max N-gram

๐Ÿ“š Model Information & Documentation

Learn more about the algorithms used in this tool:



This Keyword Extraction Explorer Tool was created as part of the Digital Scholarship at Oxford (DiSc) funded research project: Extracting Keywords from Crowdsourced Collections.

See also the main EKCC project repository: Crowdsourced Data Tools โ†—

and Arana-Catania, M., Conisbee, C., & Kidd, M. (2026). Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI. https://doi.org/10.48550/arXiv.2607.09324 โ†—

The code for this tool was built with the aid of Claude Opus 4 and updated with Opus 4.8.