Monday, 9 November 2009

63613 words

and most of them in the right order.

today's error

I am losing the will to live! All I want to do is write, but I keep finding minor errors in my data. So frustrating, but hopefully I am catching all the errors.

Today's blooper. In the distance to region centre I got the centroid wrong for Sheffield City Centre (fat fingers), thus all the refs were about 200 miles away in a big hump on the graph. Great there is a hump, but why over there? Duly corrected and now must run 4 sets of data off, then accumulate it then run the reports perform all the changes to get an excel chart that can actually be read (esp in b&w), and put it in the document. Then I can start thinking about what it means.

At least there are only 4 vernacular regions not 200 like with the administrative regions, so it is quicker to correct.

Rob.

Friday, 30 October 2009

error!

I have been sorting out the graphs for the document. not as easy as it could be using excel and knowing it will have to be b&w for the thesis. Anyway as I have been outputting the results, they have looked different to before with lovely smooth curves from the web and knobbly ones from the corpus, also far fewer georefs in the corpus. I was going to output all the graphs and then try to understand what had happened.

Turns out I got them round the wrong way, an early error and the fact that I reference the data sets by a code rater than anything descriptive meant I have been disproving my theory all week. When I saw the smooth graphs I could not beleive they were my corpus ones because that would suggest that my corpus is better not just the same as the web. I also used the wrong corpus set, I used first 100 docs whereas I have another set that matches the file quantities for the web crawl. The correct set is bulding now, it is taking an age because as I first though it has 5x (about 5 million) the georefs in it. Someone recently called me "the stupidest clever person she knows"; Mmmmm.

I think the results will be quite good, when I get them.

I am going to have to work the weekend mostly, because I spent more time that I should have going to meetings that did not happen because the person who called them did not turn up themselves. If they could only have told me I would have saved 5 hours of my time in travel and sitting about (no fun when you are not paid bth and are already working every Saturday). Oh, and apparently I didn't need to be there anyway.

Wednesday, 14 October 2009

issues

There were some issues with the last post. The number of pages for each region name from ech source was different, and I counted total number, not mean per file. Now created a new set from the corpus tht mirrors the numbers of files from web. The index is still different because I use Lucene, and who knows how Yahoo! do it. Thus even when a region name is in the settlement set, the pages retreived by my index differs from the web one.

Tuesday, 13 October 2009

ambiguity and frequency

Strange. There are many more unique references in the geocode of the corpus comparedto straight from the web, but as a raw count the difference is not nearly do pronounced. Must now check distance to centroid and other such stats.

Thursday, 17 September 2009

Don't follow the GPS too slavishly!

http://www.telegraph.co.uk/motoring/news/6197826/Driver-followed-satnav-to-edge-of-100ft-drop.html

Wednesday, 16 September 2009

Head down writing

48,485

I've been a bit quiet here.

I had a few family problems, but those are improving. They occupied me whilst I was supposed to be on holiday, so only 1/2 a holiday.

I am producing results, and statistics that allow me to assess them. I am also looking at the theory behind KDE surfaces; a bit mind boggling.