Natural Language Processing (NLP)
Speech recognition, speech translation, understanding complete sentences, understanding synonyms of matching words, sentiment analysis, and writing complete grammatically correct sentences and paragraphs.
- Outline of natural language processing | Wikipedia
- Natural Language Processing | Wikipedia
- Grammar Induction | Wikipedia
- The Stanford Natural Language Inference (SNLI) Corpus
- How do I learn Natural Language Processing? | Sanket Gupta
- AI-Powered Search
- Natural Language | Chris Umbel
- NLP News | Sebastian Ruder
- Language services | Cognitive Services | Microsoft Azure
- Text Transfer Learning
Contents
- 1 Pipeline
- 2 Regular Expressions (Regex)
- 3 Tokenization / Sentence Splitting
- 4 Stop Words
- 5 Stemming (Morphological Similarity)
- 6 Part-of-Speech (POS) Tagging
- 7 Chunking
- 8 Chinking
- 9 Named Entity Recognition (NER)
- 10 Coreference
- 11 Hierarchical Classifier
- 12 Lemmatization
- 13 Corpora
- 14 Topic Modeling
- 15 Word Embeddings
- 16 Summarizer
- 17 Ontologies
- 18 Natural Language Inference (NLI) and Recognizing Textual Entailment (RTE)
- 19 Deep Learning Algorithms
- 20 Evaluation Measures - Classification Performance
- 21 Capabilities
Pipeline
- Natural Language Tools:
- Outline of natural language processing | Wikipedia
- CogComp NLP Pipeline | Cognitive Computation Group, led by Prof. Dan Roth
Regular Expressions (Regex)
Search for text patterns, validate emails and URLs, capture information, and use patterns to save development time.
Tokenization / Sentence Splitting
Stop Words
Stemming (Morphological Similarity)
Refers to a crude heuristic process that chops off the ends of words in the hope of achieving this goal correctly most of the time, and often includes the removal of derivational affixes.
Part-of-Speech (POS) Tagging
Chunking
Chinking
Named Entity Recognition (NER)
Coreference
Hierarchical Classifier
Lemmatization
Lemmatization usually refers to doing things properly with the use of a vocabulary and morphological analysis of words, normally aiming to remove inflectional endings only and to return the base or dictionary form of a word, which is known as the lemma . If confronted with the token saw, stemming might return just s, whereas lemmatization would attempt to return either see or saw depending on whether the use of the token was as a verb or a noun. The two may also differ in that stemming most commonly collapses derivationally related words, whereas lemmatization commonly only collapses the different inflectional forms of a lemma. Stemming and lemmatization | Stanford.edu
Corpora
Topic Modeling
Word Embeddings
Summarizer
Ontologies
(aka knowledge graph) can incorporate computable descriptions that can bring insight in a wide set of compelling applications including more precise knowledge capture, semantic data integration, sophisticated query answering, and powerful association mining - thereby delivering key value for health care and the life sciences.
Natural Language Inference (NLI) and Recognizing Textual Entailment (RTE)
The goal of identifying textual entailment – whether one piece of text can be plausibly inferred from another. The current work exhibits strong ties to some earlier lines of research, particularly automatic acquisition of paraphrases and lexical semantic relationships and unsupervised inference in applications such as question answering, information extraction and summarization. It has also opened the way to newer lines of research on more involved inference methods, on knowledge representations needed to support this natural language understanding challenge and on the use of learning methods in this context. Recognizing textual entailment: Rational, evaluation and approaches | Ido Dagan, Bill Dolan, Bernrdo Magnini and Dan Roth
Semantic Role Labeling (SRL)
identifies shallow semantic information in a given sentence. The tool labels verb-argument structure, identifying who did what to whom by assigning roles that indicate the agent, patient, and theme of each verb to constituents of the sentence representing entities related by the verb.
Deep Learning Algorithms
- Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and Recurrent Neural Network (RNN)
- Deep Q Learning (DQN)
- 7 types of Artificial Neural Networks for Natural Language Processing
- Combination of Convolutional and Recurrent Neural Network for Sentiment Analysis of Short Texts | Xingyou Wang, Weijie Jiang, Zhiyong Luo
- Autoencoders / Encoder-Decoders
- LSTM
- GRU
Evaluation Measures - Classification Performance
Confusion Matrix, Precision, Recall, F Score, ROC Curves, trade off between True Positive Rate and False Positive Rate.
Capabilities
Sentiment Analysis
Wikifier