Natural Language Processing for Innovating Behavioral Political Science Research

The Oxford Handbook of Behavioral Political Science ◽

10.1093/oxfordhb/9780190634131.013.33 ◽

2020 ◽

Author(s):

Quan Li

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Political Science ◽

Language Processing ◽

Analytical Procedure ◽

Science Research ◽

Text Data ◽

Political Actors ◽

The Social ◽

Behavioral Mechanisms

Since the invention of Word2Vec by a Google team in 2013, natural language processing (NLP) techniques have been increasingly applied in the private sector, by government agencies across countries, and in the social sciences. This chapter explains NLP’s basic analytical procedure from preprocessing of raw text data to statistical modeling, reviews the most recent advances in NLP applications in political science, and proposes a new scaling approach for measuring political actors’ spatial preferences along with potential application in decision-making research. It argues that with a greater focus on explaining behavioral mechanisms and processes, which is a goal shared by artificial intelligence/computational modeling and cognitive science, NLP can help improve behavioral political science by its ability to integrate micro-, meso-, and macro-level analyses. Critical and reflexive use of NLP techniques, combined with big data, will lead to obtain better insights on political behavior in general.

Download Full-text

How Language Shapes Prejudice Against Women: An Examination Across 45 World Languages

10.31234/osf.io/mrbcf ◽

2020 ◽

Author(s):

David DeFranza ◽

Himanshu Mishra ◽

Arul Mishra

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Ongoing Debate ◽

Text Data ◽

Gender Prejudice ◽

World Languages ◽

The World ◽

Present Context ◽

The Common

Language provides an ever-present context for our cognitions and has the ability to shape them. Languages across the world can be gendered (language in which the form of noun, verb, or pronoun is presented as female or male) versus genderless. In an ongoing debate, one stream of research suggests that gendered languages are more likely to display gender prejudice than genderless languages. However, another stream of research suggests that language does not have the ability to shape gender prejudice. In this research, we contribute to the debate by using a Natural Language Processing (NLP) method which captures the meaning of a word from the context in which it occurs. Using text data from Wikipedia and the Common Crawl project (which contains text from billions of publicly facing websites) across 45 world languages, covering the majority of the world’s population, we test for gender prejudice in gendered and genderless languages. We find that gender prejudice occurs more in gendered rather than genderless languages. Moreover, we examine whether genderedness of language influences the stereotypic dimensions of warmth and competence utilizing the same NLP method.

Download Full-text

Sentiment of App with Word Vectors

International Journal of Engineering and Advanced Technology - Regular Issue ◽

10.35940/ijeat.f1416.0986s319 ◽

2019 ◽

Vol 8 (6S3) ◽

pp. 2156-2159

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Sentiment Analysis ◽

Language Processing ◽

Text Data ◽

Vector Representations ◽

Text Sentiment Analysis

Vector representations for language have been shown to be useful in a number of Natural Language Processing tasks. In this paper, we aim to investigate the effectiveness of word vector representations for the problem of Sentiment Analysis. In particular, we target three sub-tasks namely sentiment words extraction, polarity of sentiment words detection, and text sentiment prediction. We investigate the effectiveness of vector representations over different text data and evaluate the quality of domain-dependent vectors. Vector representations has been used to compute various vector-based features and conduct systematically experiments to demonstrate their effectiveness. Using simple vector based features can achieve better results for text sentiment analysis of APP.

Download Full-text

Detecting Fake News Using Deep Learning and NLP

Advances in Digital Crime, Forensics, and Cyber Terrorism - Confluence of AI, Machine, and Deep Learning in Cyber Forensics ◽

10.4018/978-1-7998-4900-1.ch007 ◽

2021 ◽

pp. 117-133

Author(s):

Uma Maheswari Sadasivam ◽

Nitin Ganesan

Keyword(s):

Deep Learning ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Reliable Information ◽

Citizen Journalism ◽

Fake News ◽

Social Unrest ◽

Safe Place ◽

The Social

Fake news is the word making more talk these days be it election, COVID 19 pandemic, or any social unrest. Many social websites have started to fact check the news or articles posted on their websites. The reason being these fake news creates confusion, chaos, misleading the community and society. In this cyber era, citizen journalism is happening more where citizens do the collection, reporting, dissemination, and analyse news or information. This means anyone can publish news on the social websites and lead to unreliable information from the readers' points of view as well. In order to make every nation or country safe place to live by holding a fair and square election, to stop spreading hatred on race, religion, caste, creed, also to have reliable information about COVID 19, and finally from any social unrest, we need to keep a tab on fake news. This chapter presents a way to detect fake news using deep learning technique and natural language processing.

Download Full-text

System for monitoring natural disasters using natural language processing in the social network Twitter

2016 IEEE International Carnahan Conference on Security Technology (ICCST) ◽

10.1109/ccst.2016.7815686 ◽

2016 ◽

Cited By ~ 2

Author(s):

Miguel Maldonado ◽

Darwin Alulema ◽

Derlin Morocho ◽

Marida Proano

Keyword(s):

Natural Language Processing ◽

Social Network ◽

Natural Language ◽

Natural Disasters ◽

Language Processing ◽

The Social

Download Full-text

IDENTIFYING BEST PRACTICES FOR USE OF TEXT DATA IN HEALTH ECONOMICS AND OUTCOMES RESEARCH USING NATURAL LANGUAGE PROCESSING

Value in Health ◽

10.1016/j.jval.2016.03.1776 ◽

2016 ◽

Vol 19 (3) ◽

pp. A82

Author(s):

B.A. Feinberg ◽

L. Lal ◽

D.F. Garofalo ◽

U. Mujumdar

Keyword(s):

Natural Language Processing ◽

Health Economics ◽

Best Practices ◽

Natural Language ◽

Language Processing ◽

Outcomes Research ◽

Text Data

Download Full-text

Natural language processing based advanced method of unnecessary video detection

International Journal of Electrical and Computer Engineering (IJECE) ◽

10.11591/ijece.v11i6.pp5411-5419 ◽

2021 ◽

Vol 11 (6) ◽

pp. 5411

Author(s):

Nazmun Nessa Moon ◽

Imrus Salehin ◽

Masuma Parvin ◽

Md. Mehedi Hasan ◽

Iftakhar Mohammad Talha ◽

...

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Detection System ◽

Video Content ◽

Text Data ◽

Data Set ◽

Plain Text ◽

Video Detection ◽

Library Function

<span>In this study we have described the process of identifying unnecessary video using an advanced combined method of natural language processing and machine learning. The system also includes a framework that contains analytics databases and which helps to find statistical accuracy and can detect, accept or reject unnecessary and unethical video content. In our video detection system, we extract text data from video content in two steps, first from video to MPEG-1 audio layer 3 (MP3) and then from MP3 to WAV format. We have used the text part of natural language processing to analyze and prepare the data set. We use both Naive Bayes and logistic regression classification algorithms in this detection system to determine the best accuracy for our system. In our research, our video MP4 data has converted to plain text data using the python advance library function. This brief study discusses the identification of unauthorized, unsocial, unnecessary, unfinished, and malicious videos when using oral video record data. By analyzing our data sets through this advanced model, we can decide which videos should be accepted or rejected for the further actions.</span>

Download Full-text

Extracting Clinical Features From Dictated Ambulatory Consult Notes Using a Commercially Available Natural Language Processing Tool: Pilot, Retrospective, Cross-Sectional Validation Study (Preprint)

10.2196/preprints.12575 ◽

2018 ◽

Author(s):

Jeremy Petch ◽

Jane Batt ◽

Joshua Murray ◽

Muhammad Mamdani

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Clinical Features ◽

Language Processing ◽

A Priori ◽

Cross Sectional Study ◽

Free Text ◽

Cross Sectional ◽

Text Data ◽

Complex Features

BACKGROUND The increasing adoption of electronic health records (EHRs) in clinical practice holds the promise of improving care and advancing research by serving as a rich source of data, but most EHRs allow clinicians to enter data in a text format without much structure. Natural language processing (NLP) may reduce reliance on manual abstraction of these text data by extracting clinical features directly from unstructured clinical digital text data and converting them into structured data. OBJECTIVE This study aimed to assess the performance of a commercially available NLP tool for extracting clinical features from free-text consult notes. METHODS We conducted a pilot, retrospective, cross-sectional study of the accuracy of NLP from dictated consult notes from our tuberculosis clinic with manual chart abstraction as the reference standard. Consult notes for 130 patients were extracted and processed using NLP. We extracted 15 clinical features from these consult notes and grouped them a priori into categories of simple, moderate, and complex for analysis. RESULTS For the primary outcome of overall accuracy, NLP performed best for features classified as simple, achieving an overall accuracy of 96% (95% CI 94.3-97.6). Performance was slightly lower for features of moderate clinical and linguistic complexity at 93% (95% CI 91.1-94.4), and lowest for complex features at 91% (95% CI 87.3-93.1). CONCLUSIONS The findings of this study support the use of NLP for extracting clinical features from dictated consult notes in the setting of a tuberculosis clinic. Further research is needed to fully establish the validity of NLP for this and other purposes.

Download Full-text

Concept of TF-IDF, Common Bag of Word and Word Embedding for Effective Sentiment Classification

International Journal of Innovative Technology and Exploring Engineering - Special Issue ◽

10.35940/ijitee.f4582.049620 ◽

2020 ◽

Vol 9 (4) ◽

pp. 2198-2201

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Sentiment Classification ◽

Word Embedding ◽

Text Representation ◽

Human Beings ◽

Text Data

Sentiment Classification is one of the well-known and most popular domain of machine learning and natural language processing. An algorithm is developed to understand the opinion of an entity similar to human beings. This research fining article presents a similar to the mention above. Concept of natural language processing is considered for text representation. Later novel word embedding model is proposed for effective classification of the data. Tf-IDF and Common BoW representation models were considered for representation of text data. Importance of these models are discussed in the respective sections. The proposed is testing using IMDB datasets. 50% training and 50% testing with three random shuffling of the datasets are used for evaluation of the model.

Download Full-text

Mining Numbers in Text: A Survey

10.5772/intechopen.98540 ◽

2021 ◽

Author(s):

Minoru Yoshida ◽

Kenji Kita

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Ad Hoc ◽

Text Data ◽

Research Areas ◽

Recent Growth ◽

Recent Advances ◽

Almost All

Both words and numerals are tokens found in almost all documents but they have different properties. However, relatively little attention has been paid in numerals found in texts and many systems treated the numbers found in the document in ad-hoc ways, such as regarded them as mere strings in the same way as words, normalized them to zeros, or simply ignored them. Recent growth of natural language processing (NLP) research areas has change this situations and more and more attentions have been paid to the numeracy in documents. In this survey, we provide a quick overview of the history and recent advances of the research of mining such relations between numerals and words found in text data.

Download Full-text

A Natural Language Processing Approach to Social License Management

Sustainability ◽

10.3390/su12208441 ◽

2020 ◽

Vol 12 (20) ◽

pp. 8441

Author(s):

Robert G. Boutilier ◽

Kyle Bahr

Keyword(s):

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Social Acceptance ◽

Rating Scale ◽

Temporal Trends ◽

Complex Projects ◽

Social License ◽

The Social ◽

Project Life

Dealing with the social and political impacts of large complex projects requires monitoring and responding to concerns from an ever-evolving network of stakeholders. This paper describes the use of text analysis algorithms to identify stakeholders’ concerns across the project life cycle. The social license (SL) concept has been used to monitor the level of social acceptance of a project. That acceptance can be assessed from the texts produced by stakeholders on sources ranging from social media to personal interviews. The same texts also contain information on the substance of stakeholders’ concerns. Until recently, extracting that information necessitated manual coding by humans, which is a method that takes too long to be useful in time-sensitive projects. Using natural language processing algorithms, we designed a program that assesses the SL level and identifies stakeholders’ concerns in a few hours. To validate the program, we compared it to human coding of interview texts from a Bolivian mining project from 2009 to 2018. The program’s estimation of the annual average SL was significantly correlated with rating scale measures. The topics of concern identified by the program matched the most mentioned categories defined by human coders and identified the same temporal trends.

Download Full-text