Rain Prediction Using Rule-Based Machine Learning Approach

Lipophilicity is a fundamental structural property that influences almost every aspect of drug discovery. Within Pfizer, we have two complementary high-throughput screens for measuring lipophilicity as a distribution coefficient (LogD) – a miniaturized shake-flask method (SFLogD) and a chromatographic method (ELogD). The results from these two assays are not the same (see Figure 1), with each assay being applicable or more reliable in particular chemical spaces. In addition to LogD assays, the ability to predict the LogD value for virtual compounds is equally vital. Here we present an in-silico LogD model, applicable to all chemical spaces, based on the integration of the LogD data from both assays. We developed two approaches towards a single LogD model – a Rule-based and a Machine Learning approach. Ultimately, the Machine Learning LogD model was found to be superior to both internally developed and commercial LogD models.<br>

Download Full-text

A Machine Learning Approach to Assess Runway Conditions Using Weather Data

10.3850/978-981-18-2016-8_474-cd ◽

2021 ◽

Author(s):

Alise Danielle Midtfjord ◽

Arne Bang Huseby

Keyword(s):

Machine Learning ◽

Weather Data ◽

Learning Approach ◽

Machine Learning Approach

Download Full-text

Redewiedergabe – Schritte zur automatischen Erkennung

Zeitschrift für Germanistische Linguistik ◽

10.1515/zgl-2019-0007 ◽

2019 ◽

Vol 47 (1) ◽

pp. 216-248

Author(s):

Annelen Brunner

Keyword(s):

Machine Learning ◽

Automatic Detection ◽

Learning Approach ◽

Manual Annotation ◽

Narrative Texts ◽

Rule Based ◽

Machine Learning Approach ◽

Different Types ◽

Short Narrative ◽

Over Time

Abstract This contribution presents a quantitative approach to speech, thought and writing representation (ST&WR) and steps towards its automatic detection. Automatic detection is necessary for studying ST&WR in a large number of texts and thus identifying developments in form and usage over time and in different types of texts. The contribution summarizes results of a pilot study: First, it describes the manual annotation of a corpus of short narrative texts in relation to linguistic descriptions of ST&WR. Then, two different techniques of automatic detection – a rule-based and a machine learning approach – are described and compared. Evaluation of the results shows success with automatic detection, especially for direct and indirect ST&WR.

Download Full-text

Teaching Machines to Find Names

Encyclopedia of Artificial Intelligence ◽

10.4018/978-1-59904-849-9.ch229 ◽

2011 ◽

pp. 1562-1567

Author(s):

Raymond Chiong

Keyword(s):

Machine Learning ◽

Statistical Models ◽

Learning Approach ◽

Named Entities ◽

Rule Based ◽

Named Entity ◽

Machine Learning Approach ◽

Wide Availability ◽

Increasing Demand ◽

Rule Based Systems

In the field of Natural Language Processing, one of the very important research areas of Information Extraction (IE) comes in Named Entity Recognition (NER). NER is a subtask of IE that seeks to identify and classify the predefined categories of named entities in text documents. Considerable amount of work has been done on NER in recent years due to the increasing demand of automated texts and the wide availability of electronic corpora. While it is relatively easy and natural for a human reader to read and understand the context of a given article, getting a machine to understand and differentiate between words is a big challenge. For instance, the word ‘brown’ may refer to a person called Mr. Brown, or the colour of an item which is brown. Human readers can easily discern the meaning of the word by looking at the context of that particular sentence, but it would be almost impossible for a computer to interpret it without any additional information. To deal with the issue, researchers in NER field have proposed various rule-based systems (Wakao, Gaizauskas & Wilks, 1996; Krupka & Hausman, 1998; Maynard, Tablan, Ursu, Cunningham & Wilks, 2001). These systems are able to achieve high accuracy in recognition with the help of some lists of known named entities called gazetteers. The problem with rule-based approach is that it lacks the robustness and portability. It incurs steep maintenance cost especially when new rules need to be introduced for some new information or new domains. A better option is thus to use machine learning approach that is trainable and adaptable. Three wellknown machine learning approaches that have been used extensively in NER are Hidden Markov Model (HMM), Maximum Entropy Model (MEM) and Decision Tree. Many of the existing machine learning-based NER systems (Bikel, Schwartz & Weischedel, 1999; Zhou & Su, 2002; Borthwick, Sterling, Agichten & Grisham, 1998; Bender, Och & Ney, 2003; Chieu & Ng, 2002; Sekine, Grisham & Shinnou, 1998) are able to achieve near-human performance for named entity tagging, even though the overall performance is still about 2% short from the rule-based systems. There have also been many attempts to improve the performance of NER using a hybrid approach with the combination of handcrafted rules and statistical models (Mikheev, Moens & Grover, 1999; Srihari & Li, 2000; Seon, Ko, Kim & Seo, 2001). These systems can achieve relatively good performance in the targeted domains owing to the comprehensive handcrafted rules. Nevertheless, the portability problem still remains unsolved when it comes to dealing with NER in various domains. As such, this article presents a hybrid machine learning approach using MEM and HMM successively. The reason for using two statistical models in succession instead of one is due to the distinctive nature of the two models. HMM is able to achieve better performance than any other statistical models, and is generally regarded as the most successful one in machine learning approach. However, it suffers from sparseness problem, which means considerable amount of data is needed for it to achieve acceptable performance. On the other hand, MEM is able to maintain reasonable performance even when there is little data available for training purpose. The idea is therefore to walkthrough the testing corpus using MEM first in order to generate a temporary tagging result, while this procedure can be simultaneously used as a training process for HMM. During the second walkthrough, the corpus uses HMM for the final tagging. In this process, the temporary tagging result generated by MEM will be used as a reference for subsequent error checking and correction. In the case when there is little training data available, the final result can still be reliable based on the contribution of the initial MEM tagging result.

Download Full-text

Providing the ‘Best’ Lipophilicity Assessment in a Drug Discovery Environment

10.26434/chemrxiv.14292485.v1 ◽

2021 ◽

Author(s):

george chang ◽

Nathaniel Woody ◽

Christopher Keefer

Keyword(s):

Machine Learning ◽

Drug Discovery ◽

High Throughput ◽

In Silico ◽

Shake Flask ◽

Chromatographic Method ◽

Learning Approach ◽

Rule Based ◽

Machine Learning Approach ◽

High Throughput Screens

Lipophilicity is a fundamental structural property that influences almost every aspect of drug discovery. Within Pfizer, we have two complementary high-throughput screens for measuring lipophilicity as a distribution coefficient (LogD) – a miniaturized shake-flask method (SFLogD) and a chromatographic method (ELogD). The results from these two assays are not the same (see Figure 1), with each assay being applicable or more reliable in particular chemical spaces. In addition to LogD assays, the ability to predict the LogD value for virtual compounds is equally vital. Here we present an in-silico LogD model, applicable to all chemical spaces, based on the integration of the LogD data from both assays. We developed two approaches towards a single LogD model – a Rule-based and a Machine Learning approach. Ultimately, the Machine Learning LogD model was found to be superior to both internally developed and commercial LogD models.<br>

Download Full-text

Outlier detection in WSN by entropy based machine learning approach

Indonesian Journal of Electrical Engineering and Computer Science ◽

10.11591/ijeecs.v20.i3.pp1435-1443 ◽

2020 ◽

Vol 20 (3) ◽

pp. 1435

Author(s):

Manmohan Singh Yadav ◽

Shish Ahamad

Keyword(s):

Machine Learning ◽

Outlier Detection ◽

Sensor Data ◽

Learning Approach ◽

Random Subspace ◽

Environmental Disasters ◽

Machine Learning Approach ◽

The World ◽

Prediction Capability

<p>Environmental disasters like flooding, earthquake etc. causes catastrophic effects all over the world. WSN based techniques have become popular in susceptibility modelling of such disaster due to their greater strength and efficiency in the prediction of such threats. This paper demonstrates the machine learning-based approach to predict outlier in sensor data with bagging, boosting, random subspace, SVM and KNN based frameworks for outlier prediction using a WSN data. First of all database is pre processed with 14 sensor motes with presence of outlier due to intrusion. Subsequently segmented database is created from sensor pairs. Finally, the data entropy is calculated and used as a feature to determine the presence of outlier used different approach. Results show that the KNN model has the highest prediction capability for outlier assessment.</p>

Download Full-text

A Machine Learning Approach to Predict Energy Consumptions in Office and Industrial Buildings as a Function of Weather Data

TECNICA ITALIANA-Italian Journal of Engineering Science ◽

10.18280/ti-ijes.632-449 ◽

2019 ◽

Vol 63 (2-4) ◽

pp. 452-458

Author(s):

Francesco Martellotta ◽

Alessandro Cannavale ◽

Ubaldo Ayr

Keyword(s):

Machine Learning ◽

Weather Data ◽

Learning Approach ◽

Industrial Buildings ◽

Machine Learning Approach ◽

Energy Consumptions

Download Full-text

A MACHINE LEARNING APPROACH WITH VERIFICATION OF PREDICTIONS AND ASSISTED SUPERVISION FOR A RULE-BASED NETWORK INTRUSION DETECTION SYSTEM

Proceedings of the Fourth International Conference on Web Information Systems and Technologies ◽

10.5220/0001524801430148 ◽

2008 ◽

Keyword(s):

Machine Learning ◽

Intrusion Detection ◽

Intrusion Detection System ◽

Detection System ◽

Network Intrusion Detection ◽

Learning Approach ◽

Rule Based ◽

Network Intrusion ◽

Network Intrusion Detection System ◽

Machine Learning Approach

Download Full-text

A Comprehensive Review on Online News Popularity Prediction using Machine Learning Approach

SMART MOVES JOURNAL IJOSCIENCE ◽

10.24113/ijoscience.v5i1.181 ◽

2019 ◽

Vol 5 (1) ◽

pp. 7

Author(s):

Priyanka Rathord ◽

Dr. Anurag Jain ◽

Chetan Agrawal

Keyword(s):

Machine Learning ◽

Social Media ◽

Online News ◽

Machine Learning Techniques ◽

News Article ◽

Learning Approach ◽

Learning Techniques ◽

Machine Learning Approach ◽

The World ◽

Popularity Prediction

With the help of Internet, the online news can be instantly spread around the world. Most of peoples now have the habit of reading and sharing news online, for instance, using social media like Twitter and Facebook. Typically, the news popularity can be indicated by the number of reads, likes or shares. For the online news stake holders such as content providers or advertisers, it’s very valuable if the popularity of the news articles can be accurately predicted prior to the publication. Thus, it is interesting and meaningful to use the machine learning techniques to predict the popularity of online news articles. Various works have been done in prediction of online news popularity. Popularity of news depends upon various features like sharing of online news on social media, comments of visitors for news, likes for news articles etc. It is necessary to know what makes one online news article more popular than another article. Unpopular articles need to get optimize for further popularity. In this paper, different methodologies are analyzed which predict the popularity of online news articles. These methodologies are compared, their parameters are considered and improvements are suggested. The proposed methodology describes online news popularity predicting system.

Download Full-text