# Link Grammar Parser: Typed Link Graphs for NLP in English, Thai, Russian, and Arabic

> opencog/link-grammar is the active continuation of the CMU Link Grammar Parser, a C-based natural language parser that produces typed link graphs rather than conventional constituency or dependency trees. Version 5.13.0 supports English, Thai, Russian, Arabic, Persian, and limited subsets of additional languages, with a multi-threaded, UTF-8 clean implementation suitable for cloud deployment.

**opencog/link-grammar** — The CMU Link Grammar natural language parser

- Repository: https://github.com/opencog/link-grammar
- Stars: 422 · Forks: 122
- Language: C
- License: LGPL-2.1
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/opencog-link-grammar

## What Link Grammar Produces and Why It Differs from Dependency Parsing

Standard NLP parsers produce either constituency trees (HPSG/phrase structure) or dependency graphs that label grammatical relations like subject, object, and modifier. Link Grammar produces neither directly. Its output is a planar graph of typed links between words in a sentence, where each link type encodes a specific grammatical relationship.

Consider a simple English sentence like "This is a test." In the Link Grammar output, the parser emits links including Wd (head-noun pointer from the wall), Ss*b (subject agreement link between the subject and verb), Ost (object link from the verb to the noun), and Xp (punctuation link). The link type Ss*b is not just a subject link: it also encodes that the subject is singular, expressed in the link's subscript notation. The type Ds**c connecting the determiner to the noun encodes both singularity and the fact that the noun begins with a consonant, distinguishing contexts where 'a' versus 'an' would be used.

The README describes this as going a bit deeper into the syntactico-semantic structure of a sentence: it provides considerably more fine-grained and detailed information than what is commonly available in conventional parsers. Constituency and dependency representations can be derived from Link Grammar output by applying conversion rules, which the README notes is possible precisely because of the richer underlying annotation.

This depth has a practical consequence. The bottom portion of the Link Grammar output lists disjuncts for each word, which are essentially very fine-grained parts of speech. The disjunct S- O+ marks a transitive verb, while the additional markup on 'is' encodes that it is specifically a transitive verb that took a singular subject and can serve as the head verb. This level of granularity is useful for grammar checking, linguistic research, and semantic reasoning, but represents significant complexity compared to simpler dependency parsers.

## Language Coverage: English, Thai, Russian, Arabic, and Persian

The parser ships with dictionaries for several languages, with coverage depth varying significantly. English is the most developed: the README notes that parse coverage of English has been dramatically improved in the current codebase compared to earlier versions.

The Thai dictionary is described in the README as now fully developed, effectively covering the entire language. Thai presents specific structural challenges because syntactic structure in Thai is carried by suffixes rather than stems. The Link Grammar parser handles this through morphological analysis: the LL link connects a Thai stem to its suffix, and the syntactic links attach only to the suffix, matching the language's actual structure.

Russian support includes morphological analysis at the word level. A Russian example in the README shows the parser splitting a word into its stem and suffix and assigning syntactic links to the suffix component.

Arabic and Persian are listed but the README does not characterize their coverage depth the way it does for Thai. The README's opening sentence describes support for English, Thai, Russian, Arabic, and Persian, and then adds "limited subsets of a half-dozen other languages," which signals that the other language dictionaries are research-grade rather than production-grade.

The parser's morphological analysis is not limited to those languages. The README describes a sophisticated tokenizer that can offer alternative splittings for morphologically ambiguous words, which generalizes across languages with productive morphology.

## Building and Using the Command-Line Parser

The repository follows a standard autotools build process. The top-level autogen.sh script generates the build system:

```bash
./autogen.sh
```

After configuration and building, the linkparser command provides an interactive mode for exploring the parser's output. Entering a sentence produces the full link diagram, including the link graph, the constituent tree derived from it, and the disjunct listing for each word.

The parser also accepts input from files and standard input for batch processing. The man page (man link-parser in the man/ subdirectory) documents the full command-line interface.

The repository includes Docker support in the docker/ directory. For cloud deployment or containerized environments, the Docker path avoids the need to manage the build toolchain locally.

As of version 5.9.0, the parser includes an experimental sentence generator. This is a fill-in-the-blanks API where words are substituted into wildcard positions whenever the result produces a grammatically valid sentence according to the Link Grammar rules. The man page for this feature is man link-generator. The generator is used in the OpenCog Language Learning project, which applies information-theoretic techniques to learn Link Grammars from corpora.

## API Bindings and Integration into Applications

Link Grammar provides C API bindings that are the foundation for all other language integrations. The bindings/ directory in the repository contains implementations for additional languages. The README mentions that the parser includes APIs in various programming languages alongside the command-line tool, without enumerating the full list in the visible documentation.

The C API is multi-threaded: dictionary updates and parsing are mutually thread-safe according to the README. This means the same parser instance can be called concurrently from multiple threads without locking the dictionary, which is relevant for server deployments that process multiple requests simultaneously.

Dictionary updates can happen at runtime without stopping the parser. The README describes this as enabling systems that perform continuous learning of grammar to also parse at the same time. This is primarily relevant for the OpenCog Language Learning use case, where grammar rules evolve as the system processes more corpora.

Word classes can be recognized with regular expressions, which allows the dictionary to handle named entities, dates, or other structured tokens without requiring an explicit entry for each one.

## Key Limitations: Language Coverage and Output Complexity

The single most important limitation is language coverage. For applications that need a parser supporting many languages at production quality, Link Grammar is not the right choice. The README is honest about this: it lists English, Thai, Russian, Arabic, and Persian as the supported languages, and characterizes the others as limited subsets. spaCy, by contrast, offers production-quality models for over twenty languages, each trained on large corpora with consistent coverage.

spaCy takes a different approach. It uses statistical models trained on annotated corpora and produces dependency parses using a standard label set (Universal Dependencies). It is faster to deploy for common use cases and easier to integrate into Python-heavy ML pipelines. Link Grammar's advantage is the depth and fine-grainedness of its annotations: the link type and disjunct information it produces is richer than a standard dependency label, which matters for tasks that require detailed grammatical analysis rather than broad coverage across many languages.

The output format itself is a limitation for teams accustomed to standard NLP toolchain conventions. Link Grammar produces its own link type vocabulary, not Universal Dependencies. Converting to UD is possible but adds a step.

## Maintenance, Theory Background, and License

The Link Grammar formalism was originally developed in 1991 by Davy Temperley, John Lafferty, and Daniel Sleator at Carnegie Mellon University. The three initial CMU publications remain the best introduction to the theory. The current opencog/link-grammar repository is a continuation of that original CMU code base, but the README describes it as profoundly different: there have been innumerable bug fixes, performance improvements of several orders of magnitude, multi-threading, UTF-8 support, and security hardening for cloud deployment.

The project is licensed under LGPL-2.1, which the README notes makes it freely available for both private and commercial use with few restrictions. LGPL allows linking against the library without requiring the calling application to be open source, provided the library itself remains under LGPL terms.

The last push to the repository was on 2026-09-22, and the current version is 5.13.0. The project receives regular updates and is maintained under the opencog organization on GitHub.

## Conclusion

Link Grammar is the right parser for NLP work that needs fine-grained syntactico-semantic annotations beyond what standard dependency parsers provide. It is not appropriate for projects that need fast out-of-the-box support for a wide range of languages: the README explicitly describes non-English language support as limited subsets. Before committing to it, confirm that the target language has sufficient dictionary coverage by testing on representative input with the linkparser command-line tool. The last push was on 2026-09-22, confirming the project is under active maintenance.

## FAQ

### What is Link Grammar and how does it differ from dependency parsing?

Link Grammar produces a planar graph of typed links between words in a sentence, where each link type encodes a specific grammatical relationship including agreement, morphological information, and syntactic role. The README describes this as going deeper into the syntactico-semantic structure than conventional dependency parsers. Dependency representations can be derived from the Link Grammar output by applying conversion rules.

### Which languages does the Link Grammar Parser support?

The README lists English, Thai, Russian, Arabic, and Persian as the supported languages, with the Thai dictionary described as fully covering the language. A half-dozen other languages are supported as limited subsets. English has the most mature coverage.

### What license does the Link Grammar Parser use?

The project is released under LGPL-2.1. The README states this makes it freely available for both private and commercial use with few restrictions, and LGPL permits linking the library into applications without requiring those applications to be open source.

## Sources

- [Issues](https://github.com/opencog/link-grammar/issues)
- [License: LGPL-2.1](https://github.com/opencog/link-grammar/blob/master/LICENSE)
- [opencog/link-grammar on GitHub](https://github.com/opencog/link-grammar)
- [README](https://github.com/opencog/link-grammar/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/opencog-link-grammar
