Tuesday, 28 April 2009
Sunday, 26 April 2009
The download of the first 250 pages for each os50k settlement name in the uk has finished. CICS are probably still trying to find their missing bandwidth, and the uni network is probably running twice as fast now (only joking the latency of each batch of 50 has been the main cause of the slowness of the program).
Done now, won't have to do it again! Now to index it and see if the results are better from this corpus (it certainly has more placenames in it judging by my indexing of the first 150 of each).
Done now, won't have to do it again! Now to index it and see if the results are better from this corpus (it certainly has more placenames in it judging by my indexing of the first 150 of each).
Wednesday, 22 April 2009
Some results

I have plotted distance against a metric that compares counts in bnc against counts in Geograph (modified from ACM GIS paper). Not sure I can see a relationship here. distance is x, the "common textness" measure is y, negaite numbers being more common text that geographic. clearly it clusters at shorter distances, and tends to be above the 0 line on the other metric.
I also checked how much "stopword" sets overlapped for the bnc derived top 1000 and the longest distance top 1000, they do not overlap more than random, although when you look at the words in each list they frequently seem plausible common english words. This suggests that the two lists do not validate each other, but should be combined. This however then leaves the problem of showing that they are valid to exclude. In particular I can hardly exclude all the furthest away places and then say "look the points cluster round the centroid better"!
I am working today on bucketing the numbers so it might be easier to see what is ging on and creating a graph showing the distribuion of distances, which I will thn run controlling for resource, region size and stopwordness. Hopefully these distributions will look different.
Tuesday, 14 April 2009
Now indexing the first 3 directories of os50kcorpus. I am still collecting another 2, but have run out of patience to see what exactly I have in the first 3, which is pages 1-150 of each settlement name. Will it be skewed by the way it was created?
The other thing that can happen with this corpus is a search for the 2500 regions, that I can geocoded against a random selection of pages; are there any differences in scopes etc? Since this process is slow, I might choose a smaller set of "regions" the middle set is uk settlements, which are bounded by rural areas anyway and seem different to snis and counties. There are also many more of them, and it seems to me there are too many. It would probably have been better to select about 100 - 500 of them, but to put more effort into finding neighbourhood data (like snis) for different cities. Probably a bit late now.
The other thing that can happen with this corpus is a search for the 2500 regions, that I can geocoded against a random selection of pages; are there any differences in scopes etc? Since this process is slow, I might choose a smaller set of "regions" the middle set is uk settlements, which are bounded by rural areas anyway and seem different to snis and counties. There are also many more of them, and it seems to me there are too many. It would probably have been better to select about 100 - 500 of them, but to put more effort into finding neighbourhood data (like snis) for different cities. Probably a bit late now.
Subscribe to:
Posts (Atom)


