Large-scale text clustering and exploration
Lingo4G is a document clustering engine that organizes large collections of text documents into understandable formats, usually clearly labeled thematic groups called clusters. The tool is designed to cluster collections of text documents, which can include data retrieved from PubMed and other databases. Before using the tool, users have to convert their documents in to formats such as JSON, PDF, word, HTML and plain text files. Users can also use some inbuilt datasets such as the PubMed Open Access and Clinicaltrials.gov reports. Users can find documents similar to the example seed document the user provides, based on keyword or semantic vector similarity. It helps users find documents with highly overlapping content, which may suggest the documents are duplicates or plagiarised copies. Users can use the Lingo4G Explorer app to experiment with topic extraction, clustering, 2D mapping, time-series analysis and duplicate detection.