Automated thesaurus creation from raw text corpora enables efficient lexical enrichment
Natural processing steps including tokenization, syntactic analysis, and attribute extraction provide a comprehensive approach
Results evaluated against psychological testing and artificial synonym methods ensure methodological rigor
Application to a wide range of corpora from 40,000 to 6 million characters offers practical insights
Includes step-by-step guidance for creating, implementing, and testing a first draft thesaurus
Summarized by Shop
Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showi