Showing posts with label random. Show all posts
Showing posts with label random. Show all posts

Monday, 30 July 2007

Things I Need to write down

Here a few things that I've read or come across that I need to get down before i forget.
  • Topic Detection is a problem of real time clustering (or near real time). Its purpose is to be able to develop clusters of stories in one pass given a small (possibly zero size) history of the cluster.
  • Topic detection and first story detection go hand in hand. In fact you can use almost identical algorithms with the only difference being you threshold for FSD.
  • It seems that the best results are obtained using mixture models. That is using several techniques in order to obtain a result.
  • The main trade offs that will occur in TDT is False Alarm Probability and Missed Probability. In order to get a good system we aim to try and minimise both. However they are generally trade offs. A system with low false alarm generally has a higher miss rate. The reason can be seen in the way they generalise.
  • Most of the clustering seems to use a vector space model, and then KNN clustering techniques.
  • A main issue that i need to look into is run time. That is ensure that the system works in an acceptable time limit, while at the same time doing its job.
  • I need to also ensure that the system is a technical RESEARCH PROJECT that has some interesting research as well as the system that is business orientated. This will become clearer after i type up the initial system approach, and include references. These will then make sure I'm heading in the right technical direction for the project.

Developing the Ideas

In the EBUS5003 lab today Rafael said that the best way to move on with the project. His suggestion was that I pick a research paper from the TDT TRECs and try to emulate their results, and then extend the work to my data set and my application.

This seems like a good idea as then i will at least have a bench mark of where to go. The paper that I will use is the paper in 'Topic Detection and Tracking: Event Based Information Organization', from CMU. It is Chapter 5 (Multi-strategy learning for Topic Detection and Tracking Yang et.al.)

The reason I choose this paper is that it covers all the areas that i will deal with, as well as having a quite detailed history of the way it has been constructed. The results and evaluation methodology is also well reported.

Overall the project i will be doing will concentrate on 3 tasks involved in TDT, all of which are described in detail. These tasks are:
  1. First Story Detection
  2. Topic Detection (Clustering)
  3. Tracking
My next plan is to obtain a copy of the TDT3 Corpus (which is already annotated as i do not want to spend too much time doing that). I've emailed Rafael about is as it seems you have to pay (allot) for it, which I don't really want to do.

I will also be obtaining a laptop from Dan tomorrow that i will be able to use for the duration of the project. This will be my working laptop where i will be able to use it at Vision Bytes. I will be going there more often from next week.

The full text of the article and the way that i will attack the problems will be up tomorrow.

Further I'm reinstalling Linux at home :-) For both EBUS COSC and Thesis, might even consider putting it on the laptop, but not sure. Think I'll keep this running Visa for the sanity at work.

Wednesday, 16 May 2007

now using vista

Well today I got Vista (Ultimate) installed and working. It wasn't too painful, and well its actually quite a nice operating system, lots of things i didn't even know existed. But yes so were now a Vista powered blog and project :)