I distinguish three kinds of models used in research.
The first kind generates unexpected predictions that can be tested and thus reveal new insights. These models are necessarily quite simple, because simplicity and elegance is a precondition for acceptance of a model that makes novel predictions. Why? Because simple models can be understood readily. This makes it possible for people to evaluate how the model's unexpected predictions come about.
The second kind are modifications of simple, established models that explain things the established models couldn't explain. An established model may fall from favour because it does not predict important new observations. Models that are modifications of existing models can save the original model, which may otherwise have been abandoned for a completely new, and perhaps immature model. The usefulness of this second kind of model lies in conserving and preserving valuable accepted wisdom.
Finally, there is a third kind of model. For my taste, all models belonging to this category should undergo refinement until they fall in either of the first two categories.....
So. Don't you agree there are only two kinds of useful models?
23 June 2006
Three types of models...
Posted by
Tobe Che Benjamin Freeman
at
11:43 pm
View Comments
17 June 2006
Taking the high throughput plunge
Scientists hesitate before embarking on a genome-wide investigation because some future shift in technology could reduce massively the time and cost of doing large scale work. But without making a start, these technologies won't come into being.
Without a genome-wide, high throughput approach an investigation could remain confined to a search beneath some lamp-lit corner of the genome, when the true key could be languishing elsewhere…
Or as Thomas Jenuwein of the Research Institute of Molecular Pathology, Vienna puts it "One can never be 100% ready... The rest will happen once the momentum is built up" (ref).
Jenuwein is referring to proposals by the International Human Epigenome Project to catalogue epigenetic features including DNA methylation and histone modification, the most well understood mechanisms underlying epigenetic influences on gene activity.
DNA methylation and histone modifications are tissue and developmentally specific. That means that they must be studied in each tissue separately, and at each developmental phase.
And neither process can currently be studied in a high throughput context. Methylation assays are accurate but slow and expensive, while large scale identification of histone marks is prone to problems with accuracy.
As with the Human Genome project, there is no alternative but to take the plunge.
Ref: Jane Qiu, Editor Nature Reviews Neuroscience in Nature May 11th 2006.
Posted by
Tobe Che Benjamin Freeman
at
8:26 pm
View Comments
31 May 2006
Solving differential equations?
Just now I wandered up to a colleague, a physicist with the bioinformatics team, lost in thought.
He was standing in the canteen line for lunch, face expressionless and oblivious to all around him.
"Solving differential equations?" I asked.
"No," he replied, calmly. "I've got to work out the model first".
Posted by
Tobe Che Benjamin Freeman
at
12:50 pm
View Comments
10 May 2006
Funding novel technologies
Australia's research universities received additional funding in this year's budget. For example, the government announced Aus$200M in new initiatives to assist small to medium sized businesses to commercialize new technologies.
Virginia Walsh, executive director of Australia's Group of Eight top research universities, welcomed the increased spending. But Walsh also highlighted the lack of funding for so-called "proof of concept", pre-commercial research investment, arguing that this lack "restricts the flow of new technology ventures".
It's a brave advocate of research that raises such a point during the relatively tight economic conditions that prevail today. But if Walsh is right, shifting investment from pre-commercial to commercial "innovation investment" might stymie the very conditions necessary for such investment to be a success.
The Group of Eight has also been outspoken about the way research budgets are allocated, giving special attention to the Research Quality Framework. Group of Eight Chair Glyn Davis pointed to the “lack of detail about the amount of funding to be distributed on the basis of RQF performance”.
I found Davis’ quote most interesting of all. When performance measures are used to determine funding levels, one should be reassured that performance would be rewarded. Otherwise, it would seem that the stick remains in place, but the carrot is nowhere to be seen.
Posted by
Tobe Che Benjamin Freeman
at
9:31 pm
View Comments
27 April 2006
Future shock
We look to advanced technologies to solve our problems, but what happens when we are presented with a solution that is cutting edge?
Talk to a researcher these days and they will tell you that the rate limiting step in research is no longer the gathering of data, but rather the interpretation of it. Gathering information is highly automated, but analyzing it remains a fairly manual process.
This problem is felt acutely in industrial research. I have some experience with two industrial research fields confronting this problem: pharmaceutical research in toxicology and financial risk management.
Both industries claim to be overwhelmed by the volumes of data that they must analyze and interpret. This has triggered interest in quantitative methods to analyze large, complicated (high dimensional) datasets.
These fields produce experimental results that are just too large to hold in your head. Methods such as Self Organising Maps, principal components analysis (PCA) and supervised machine learning algorithms are ideal for analyzing such data. They have been around for years, but fast, inexpensive computers and user friendly interfaces have made them available on everyone's desktop.
PCA can be used to create a view of complex (high dimensional) data based on a few dimensions. Supervised machine learning can fish out patterns in data that exist in high dimensional spaces; patterns that have a complexity that our mind cannot grasp.
So you might expect that industry would jump at the chance to use these methods. Well, not so fast.
Before I describe my experience of industry's reaction to such methods, it's worth looking at a bit of background on the life of a company toxicologist or financial risk manager.
Toxicologists and financial risk managers don't have it easy. Suppose that a medicine produces a serious unexpected side effect. All drugs have been very carefully tested in development to reduce the chance of this happening. The buck stops with the toxicologist that performs these tests.
In finance there is a similar problem. Fund managers invest money using pre-agreed strategies. The strategies have been assessed in terms of their profitability and their risk. Financial risk managers must answer for financial losses caused by events that were not anticipated in their risk estimates.
Within this context, it should not come as a surprise that these industries do not exactly leap at the chance to use PCA and machine learning algorithms. I would go as far as to say that there can be a disconnect between the eager mathematical physicist pitching a statistical method and the industrial practitioners that are their target audience.
I suspect that the problem goes deeper than mere suspicion of a new and relatively untested approach. Arguably, it lies with the very nature of the methods themselves.
Consider this. Most data analysis begins with a hunch about what the data will ultimately show. This hunch might be a correlation between two known variables, or perhaps a simple pattern of results across two or three experimental conditions. We can create a picture of these results in our head.
PCA and machine learning algorithms don't work this way. They produce projections from high dimensional spaces onto low dimensional spaces.
It is beyond our minds capacity to visualize the exact combination of the original variables that is captured by a principle component. The low dimensional projection is delivered to us without a name. The same goes for patterns derived by machine learning.
Without names, these results don't tell a story.
Skillful interpretation of a PCA can yield a story of sorts. But the power of these methods is that they create an unbiased view of the data, one that doesn't need to adhere to a pre-existing story.
I don't know about you but I find this distinction rather deep. And if you compare these approaches with more familiar analytical approaches - ones that involve hunches, stories and easy intuitions for the patterns of results, then its not surprising that my friends in industry tend to shy away from them.
But analytically, this unbiased view is a big, big plus. These are methods that can reveal unexpected patterns in results and they scale well with the size of the data set.
That is what excites the mathematical physicists.
Posted by
Tobe Che Benjamin Freeman
at
9:17 am
View Comments
Collaboration between Biopolis and RIKEN
My eyes are on Biopolis, Singapore's biosciences research initiative occupying a futuristic campus next to the National University of Singapore. Championed by Philip Yeo, Chair of Singapore's Agency for Science, Technology and Research (A*Start), Biopolis is still relatively early phase. But it already includes a crop of several shiny buildings connected by soaring above-ground walkways.
Biopolis has been making impressive connections lately, and actively working itself into the public imagination. Recent news describes a collaborative agreement with Japan's RIKEN (translated as "Institute of Physical and Chemical Research", but these days doing plenty of top biological research).
Biopolis goes into the RIKEN collaboration with an interest to expand its biomedical research focus beyond infectious disease to cancer drug development. The collaboration will focus on the exchange of ideas as well as training programs that will send Singapore's burgeoning supply of enthusiastic trainee scientists to Japan.
Biopolis
RIKEN
Posted by
Tobe Che Benjamin Freeman
at
9:05 am
View Comments
22 April 2006
Reporting research news
In journalism one must report Who, What, When, Where, Why, and How. Journalists have methods to obtain this information, and I am interested in tailoring these methods to reporting research news.
A typical reader of research news expects high quality background information surrounding the news. This information gives the reader a crash course on scientific aspects of the news topic.
This raises the reporting challenge. The journalist must describe what, why and how in depth, but can not hope to have first hand knowledge in all cases.
The journalist therefore needs a way to capture expert knowledge to include as background. I'm interested in refining methods of capturing expert knowledge.
To a professional journalist, the methods will probably end up looking like "good old-fashioned journalism". But please humour me while I find out how to give researchers a good hearing in the Press.
Posted by
Tobe Che Benjamin Freeman
at
10:18 am
View Comments
16 April 2006
mission statement
I am interested in research and how it can be fostered. Virtually all
kinds of research interest me and I believe research should be defined
broadly to include any activity that creates knowledge.
Why should research be fostered? It's difficult to fund research,
train people in research methods and to do research.
But worthwhile.
Researchers find things out for us. They gather information with which
to make sound decisions. They discover new things. They refine the
methods that make these discoveries possible. If we lost our
researchers we would make less informed decisions, fail to discover
new things and would lack the tools to train new people to replace
them.
My aim is to raise awareness of research efforts and fuel debate about
how these efforts can be refined and improved.
Posted by
Tobe Che Benjamin Freeman
at
9:56 pm
View Comments
