Posts mit dem Label science related werden angezeigt. Alle Posts anzeigen
Posts mit dem Label science related werden angezeigt. Alle Posts anzeigen

Mittwoch, 30. Juli 2014

Integrating Machine Learning Models for Real-Time Prediction into your Existing Workflow (using openscoring and PMML)

In today's world, understanding customers and learning from their behavior is a key component in a company's competitive edge in the market. This not only refers to lower user-retention costs in marketing through intelligently timed re-engagement and a higher lifetime value of users through clever item recommendation systems, but also extends to lower operating costs and a better user experience through modern risk and fraud management: discovering fraudulent participants in marketplaces or payments at risk are vital to the overall performance.

Especially for established companies outside the mobile- and web-service space, adopting new practices and integrating lessons learned from a thorough data analysis can be hard. Established work flows and running systems need to be changed, which can often be a painful experience---especially when outside consultants are hired to conduct the initial study and viability analysis.  Their tools might not be a good fit for the company's established stack.

A good solution to integrate a new layer of data mining and machine learning models is a middleware layer, such as openscoring. It runs independently and can be accessed through a REST API it provides. Existing software does not need to extended with new libraries for data analysis, but only needs to be able to communicate via HTTP, passing on XML requests---a very low bar most systems will pass without any additions---and communication that can be implemented without dedicated machine learning specialists.

The machine learning models can be created offline, in a variety of languages such as Python or R. openscoring. A XML description language, the predictive model markup language (PMML), is then used to deploy the model on a server in the cloud. There is even a heroku ready version that can be set up with a couple of lines of code in a matter of minutes.

In the following, I will outline how a model created using the statistical language R, e.g. by a group of consultants or a in-house team, is deployed as a service, ready to be integrated in your existing frameworks.

The example is very simple: Linear regression, i.e. fitting a linear function to given data points such that a given error function is minimized. You can just open your RStudio or any other environment for R and try it out yourself.

The package for R is called pmml and can be installed using the command
> install.packages("pmml")
There is good documentation available online. Since the native output of the package is XML, make sure you have the XML library installed.

The following code snipped creates a linear regression model for a data file on my web server; please beware that it omits a number of important steps (like dealing with missing data, or normalizing the data). But it suffices to give a rough idea: A model is created and fed to the pmml-function, which in turn creates a XML description. We store the description in a file named glm-pmml.xml.
library(pmml)
library(XML)

rawDataDF <- read.CSV("http://rattle.togaware.com/audit.csv")
rawDataDF <- na.omit(rawDataDF)

target <- rawDataDF$TARGET_Adjusted

N <- length(target)
M <- N-500

data.trainingIndex <- sample(N,M)
data.trainingSet <- rawDataDF[data.trainingIndex,]
data.testSet <- rawDataDF[-data.trainingIndex,]

glm.model <- glm(data.trainingSet$TARGET_Adjusted ~ ., data=data.trainingSet, family="binomial")
glm.pmml <- pmml(glm.model, name="GLM Model", data=data.trainingSet)

xmlFile <- file.path(getwd(),"glm-pmml.xml")
saveXML(glm.pmml,xmlFile)

After creating the model and storing it in a PMML file, the next step is its deployment. There are two choices: 1) the model can be uploaded via the REST interface, or 2) it can be given as a command line parameter.
1) The request using the command line tool curl just PUTs

> curl -X PUT --data-binary @glm-pmml.xml -H "Content-type: text/xml" http://localhost:8080/openscoring/model/GLMTest

2) Via command line

> java -cp client-executable-1.1-SNAPSHOT.jar org.openscoring.client.Deployer --model http://localhost:8080/openscoring/model/GLMTest --file glm-pmml.xml

The setup of openscoring itself is straightforward and uncomplicated. Either clone the git on github and deploy it directly to heroku, or download and install it locally---the documentation of openscoring as well as Maven provides step-by-step instructions.

Using a running instance of openscoring with its model is simple: Just send requests via HTTP. For the sake of simplicity, we can just feed back the whole CSV file we used for training:

> curl -X POST --data-binary @simple_model.csv -H "Content-type: text/plain" http://localhost:8080/openscoring/model/GLMTest/csv


The answer will be a list of input-output values. Instead of using curl to send requests via the command line you can easily integrate the API with your existing software projects, e.g. to receive a score to evaluate the likelihood of fraudulent offers on your marketplace. A prominent user of openscoring is AirBnb: The young company uses decision tree models employed on openscoring servers to evaluate and catch fraudulent bookings in real-time.

Are there any drawbacks to this approach? Yes, in some cases: since the machine learning models need to be supported in the PMML language, the newest ideas presented in research papers cannot directly moved into production with openscoring and PMML. But for vast majority of use cases, this certainly does not matter a lot: while new models often have a slight edge in their specific application area as presented in papers, the transfer to a company's application will not automatically translate into the same performance advantage over traditional models. The amount of fine tuning necessary to have any advantage will outbalance any disadvantage a slightly older machine learning model will have.

Stay tuned for my follow-up article covering the use of PMML for data crunching using Hadoop.


Dienstag, 3. Januar 2012

On Human Brain Size; the Conciousness and Anaesthesia

A happy new year to all readers. I just got two articles with some references on further resources:


In near future, I will post some more math; I just haven't read all the things I stumbled on.

edit: let me squeeze in another link to an article:

Language learning is fast when words are connected with movements/gestures, which is also true for abstract words with no obvious gesture for it.

Samstag, 24. Dezember 2011

On (implementation issues of) Virtual Stock Markets

After my recent post on prediction markets, I decided to implement a small auctioneer for virtual stock markets with two parties making deals, i.e. not a market maker mechanism but a (traditional) double auction. But, despite the simple idea behind a bid/ask driven market, the implementation has to take into account quite a number of details: can you buy/sell one unit or multiple units of a good at a time? Does trading occur at discrete events or is it continuous? Is pricing uniform (all trade at the same equilibirum price) or discriminatory (individual matching of bid/ask orders)? And, if one the two options is chosen, how is the price determined?

Sonntag, 27. November 2011

Nuclear Energy: Thorium Reactors

I recently discovered a small gap in my education in nuclear energy. It concerns LFTR reactors. They seem to be reconsidered after uranium based reactors were favored after world war 2 and the during the cold war due to their dual use in energy production as well as the generation of byproducs for nuclear weapons.

There are several videos on the google tech talk channlel, just two of those:



Freitag, 18. November 2011

Contemporary Physics

Just some physics stuff from recent news: An article at wired about recent findings at lhc/cern regarding the charge party violation and a confirmation about the (too) fast neurino measurements at the lhc.

Samstag, 22. Mai 2010

Recent readings

A few remarks concerning some books I read since my last post:

  • Patterns of Democracy: Government Forms and Performance in Thirty-Six Countries presents a board study of democratic governments and classifies them into 'majoritarian' (e.g. UK) and 'consensus' (e.g. Switzerland) systems. He finds 10 items in 2 dimensions (or: the 10 items can be factored in 2 factors which good correlations) which are useful for classification. 36 countries are analysed with respect to this 10 items and the classical prejudices (majoritarian governments can act quicker and thus are better and more decisive, etc) are examined. His conclusion is that most of these turn out to be false; for some, rather the opposite seems to be true.
  • Qed: The Strange Theory of Light and Matter (Princeton Science Library) is a very good popular science book - in my opinion even better for those who have studied the matter beforehand. It does not contain any formulae - something I would complain about a lot if I didn't already know many of them. Otherwise, if you haven't had contact with quantum mechanics at all, it might be too less to be useful.
  • The Elegant Universe: Superstrings, Hidden Dimensions, and the Quest for the Ultimate Theory is one of the very few books on string theory I got my hands on yet. For a math/science orientated person, it gives a very rough but imprecise picture on the possible solutions string theory /might/ offer/s. On the other hand, it skips quite some of the problems string theory actually has (like the validity of the use of pertubation theory) and mostly due to the nature of string theory itself, doesn't offer any formulae. The recapitulation of relativity and quantum mechanics in the beginning of the book is ok, but not the best I read.
  • The other one seems to be available in German only: a collection of papers on the topics of globalization and nongovernment organizations: Die Privatisierung der Weltpolitik. It sheds light on the increasing influence of non-governmental organizations in world politics - not necessarily lobbyists from economy and industry, but also public interest organizations like environmental or animal rights activists. Very enjoyable is the neutral attitude, more analyzing than taking any side.
I ordered a bunch of new books on epigenetics and neurosciences.

Sonntag, 14. März 2010

Writing systems and non-linearity

Some time ago I posted some links, one of them pointing to a non-linear, graph like writing system called Ouwi; meanwhile I spent some time reading up on the topic of writing systems and have some more pointers:

Pinuyo is a pictorial language which indicates grammatical function of the ideograms by placement and surrounding symbols.

I also found a long, but sometimes deviating thread on 2-dimensional writing systems on the conlang mailing list archive. Unfortunatelly many of the links are now dead, the thread is from 2005.

Omniglot is a nice website about alphabets, writing systems and languages. It contains some constructed systems, like block script, a sylabic alphabet for the english language which combines letters to blocks (like the korean script for example). Most of them are linear ones.

I'm still wondering if there are some languages/alphabets optimized for reading speed. Most are either evolved, natural languages, which certainly are a compromise between reading and writing and the constructed ones optimized for speed like the morse code or shorthand.

In an edit, let me just add a few more links, somewhat related:

Láadan, a feminist language designed for countering male-centering of natural languages.
The gripping language, not spoken but transmitted by touching of the hands.

Sonntag, 3. August 2008

Great pictures of the Large Hadron Collider

See

http://www.boston.com/bigpicture/2008/08/the_large_hadron_collider.html

for some great pictures. Enjoy!