Natural Language Processing (NLP)
- Outline of natural language processing | Wikipedia
- Natural Language Processing | Wikipedia
- Grammar Induction | Wikipedia
- How do I learn Natural Language Processing? | Sanket Gupta
- AI-Powered Search
- Natural Language | Chris Umbel
- NLP News | Sebastian Ruder
- Language services | Cognitive Services | Microsoft Azure
- Text Transfer Learning
Speech recognition, speech translation, understanding complete sentences, understanding synonyms of matching words, sentiment analysis, and writing complete grammatically correct sentences and paragraphs.
Contents
- 1 Regular Expressions (Regex)
- 2 Tokenization / Sentence Splitting
- 3 Stop Words
- 4 Stemming (Morphological Similarity)
- 5 Part-of-Speech (POS) Tagging
- 6 Chunking
- 7 Chinking
- 8 Named Entity Recognition (NER)
- 9 Coreference
- 10 Hierarchical Classifier
- 11 Lemmatization
- 12 Corpora
- 13 Topic Modeling
- 14 Word Embeddings
- 15 Summarizer
- 16 Ontologies
- 17 Semantic Role Labeling (SRL)
- 18 Deep Learning Algorithms
- 19 Pipeline
- 20 Evaluation Measures - Classification Performance
- 21 Capabilities
Regular Expressions (Regex)
Search for text patterns, validate emails and URLs, capture information, and use patterns to save development time.
Tokenization / Sentence Splitting
Stop Words
Stemming (Morphological Similarity)
Refers to a crude heuristic process that chops off the ends of words in the hope of achieving this goal correctly most of the time, and often includes the removal of derivational affixes.
Part-of-Speech (POS) Tagging
Chunking
Chinking
Named Entity Recognition (NER)
Coreference
Hierarchical Classifier
Lemmatization
Lemmatization usually refers to doing things properly with the use of a vocabulary and morphological analysis of words, normally aiming to remove inflectional endings only and to return the base or dictionary form of a word, which is known as the lemma . If confronted with the token saw, stemming might return just s, whereas lemmatization would attempt to return either see or saw depending on whether the use of the token was as a verb or a noun. The two may also differ in that stemming most commonly collapses derivationally related words, whereas lemmatization commonly only collapses the different inflectional forms of a lemma. Stemming and lemmatization | Stanford.edu
Corpora
Topic Modeling
Word Embeddings
Summarizer
Ontologies
(aka knowledge graph) can incorporate computable descriptions that can bring insight in a wide set of compelling applications including more precise knowledge capture, semantic data integration, sophisticated query answering, and powerful association mining - thereby delivering key value for health care and the life sciences.
Semantic Role Labeling (SRL)
identifies shallow semantic information in a given sentence. The tool labels verb-argument structure, identifying who did what to whom by assigning roles that indicate the agent, patient, and theme of each verb to constituents of the sentence representing entities related by the verb.
Deep Learning Algorithms
- Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and Recurrent Neural Network (RNN)
- Deep Q Learning (DQN)
- 7 types of Artificial Neural Networks for Natural Language Processing
- Combination of Convolutional and Recurrent Neural Network for Sentiment Analysis of Short Texts | Xingyou Wang, Weijie Jiang, Zhiyong Luo
- Autoencoders / Encoder-Decoders
- LSTM
- GRU
Pipeline
Evaluation Measures - Classification Performance
Confusion Matrix, Precision, Recall, F Score, ROC Curves, trade off between True Positive Rate and False Positive Rate.
Capabilities
Wikifier