Wednesday, 20 January 2010
Slow but steady
This document is slowly starting to look like a PhD. It was supposed to do that before Christmas, but some personal things got in the way. I handed in a draft anyway, which was a vast improvement on previous attempts, but not finished in any way really. Anyway, I now seem to be adding about 600-1000 words per day that explain things a bit more clearly, as per comments from my supervisor. It still looks a bit like 4 separate experiments, which to some extent it is, but I will work on showing how these are linked (which they are).
Monday, 11 January 2010
New year
Well it's been a busy new year. No drinkies on NYE as Maggie got a virus. I got it the following day. The effects are just about gone for both of us.
I've had to go down to Kent and discharge my Mum. She seems fine now, but we will see.
It has been snowing a bit. Not really caused me any problems as most of the time I work from home, and the internet has stayed up all the time.
I have completed my marking. The electronic hand in has been interesting. At first the office printed it, the lead tutor photocopied it and put it on my desk (now in B&w). When I got it home I found there were two copies of one assignment (it transpired that 5 had been handed in by that group). There was also one missing. In the ones that appeared ok to start with some of the diagrams did not print. I got electronic copies and one still has diagrams missing, I'm not sure if this is the offices fault or the students. It has taken more of my time than it should have to sort all this out. Still at least I'm paid for it.
I am awaiting full comments from my supervisors, but already have some from my first supervisor. I can fully familiarise myself with what I have so far, and try to get a bit more polish on it. As 1st supervisor points out, many bit of the thesis are not finished.
Rob.
Monday, 9 November 2009
today's error
I am losing the will to live! All I want to do is write, but I keep finding minor errors in my data. So frustrating, but hopefully I am catching all the errors.
Today's blooper. In the distance to region centre I got the centroid wrong for Sheffield City Centre (fat fingers), thus all the refs were about 200 miles away in a big hump on the graph. Great there is a hump, but why over there? Duly corrected and now must run 4 sets of data off, then accumulate it then run the reports perform all the changes to get an excel chart that can actually be read (esp in b&w), and put it in the document. Then I can start thinking about what it means.
At least there are only 4 vernacular regions not 200 like with the administrative regions, so it is quicker to correct.
Rob.
Friday, 30 October 2009
error!
I have been sorting out the graphs for the document. not as easy as it could be using excel and knowing it will have to be b&w for the thesis. Anyway as I have been outputting the results, they have looked different to before with lovely smooth curves from the web and knobbly ones from the corpus, also far fewer georefs in the corpus. I was going to output all the graphs and then try to understand what had happened.
Turns out I got them round the wrong way, an early error and the fact that I reference the data sets by a code rater than anything descriptive meant I have been disproving my theory all week. When I saw the smooth graphs I could not beleive they were my corpus ones because that would suggest that my corpus is better not just the same as the web. I also used the wrong corpus set, I used first 100 docs whereas I have another set that matches the file quantities for the web crawl. The correct set is bulding now, it is taking an age because as I first though it has 5x (about 5 million) the georefs in it. Someone recently called me "the stupidest clever person she knows"; Mmmmm.
I think the results will be quite good, when I get them.
I am going to have to work the weekend mostly, because I spent more time that I should have going to meetings that did not happen because the person who called them did not turn up themselves. If they could only have told me I would have saved 5 hours of my time in travel and sitting about (no fun when you are not paid bth and are already working every Saturday). Oh, and apparently I didn't need to be there anyway.
Turns out I got them round the wrong way, an early error and the fact that I reference the data sets by a code rater than anything descriptive meant I have been disproving my theory all week. When I saw the smooth graphs I could not beleive they were my corpus ones because that would suggest that my corpus is better not just the same as the web. I also used the wrong corpus set, I used first 100 docs whereas I have another set that matches the file quantities for the web crawl. The correct set is bulding now, it is taking an age because as I first though it has 5x (about 5 million) the georefs in it. Someone recently called me "the stupidest clever person she knows"; Mmmmm.
I think the results will be quite good, when I get them.
I am going to have to work the weekend mostly, because I spent more time that I should have going to meetings that did not happen because the person who called them did not turn up themselves. If they could only have told me I would have saved 5 hours of my time in travel and sitting about (no fun when you are not paid bth and are already working every Saturday). Oh, and apparently I didn't need to be there anyway.
Wednesday, 14 October 2009
issues
There were some issues with the last post. The number of pages for each region name from ech source was different, and I counted total number, not mean per file. Now created a new set from the corpus tht mirrors the numbers of files from web. The index is still different because I use Lucene, and who knows how Yahoo! do it. Thus even when a region name is in the settlement set, the pages retreived by my index differs from the web one.
Tuesday, 13 October 2009
ambiguity and frequency
Strange. There are many more unique references in the geocode of the corpus comparedto straight from the web, but as a raw count the difference is not nearly do pronounced. Must now check distance to centroid and other such stats.
Subscribe to:
Posts (Atom)