Monday, October 26, 2009

Conference: Online Humanities Scholarship

Online Humanities Scholarship: The Shape of Things to Come "is a three day conference (March 26-8, 2010) to explore how to develop and sustain online humanities research and publication. Nine scholarly papers and eighteen responses will leverage discussion by a broad group of persons invited to the conference to contribute their expertise. This group includes scholars working on other projects and persons from funding agencies, publishers, museums, libraries, and professional organizations. The conference is closed to this group in order to provide maximum focus to the discussions."

This looks to be very interesting indeed. Have a peek at resources and participants. Papers and responses are to be posted well in advance of the meeting itself. Certainly something to keep track of.

Sunday, October 25, 2009

Total Perspective Vortex

Thinking about building a renvois navigation scheme, with some kind of visualization, for the Encyclopédie, reminded me of the Total Perspective Vortex from the Hitchhiker's Guide to the Galaxy, the greatest selling electronic book in the history of the universe. It is important to note that "in an infinite universe, the one thing sentient life cannot afford to have is a sense of proportion." Thankfully, the renvois system is finite, so we won't risk brain vaporization. The original radio broadcast is available in bits and pieces on YouTube, with Don't Panic in large, friendly letters as the video track. :-) The Guide's best advice is, aside from Don't Panic, "expect the unexpected".

Wednesday, October 21, 2009

Arbre généalogique: Static Image

We periodically get requests for a high resolution image of the splendid representation of the organization of knowledge in the Encyclopédie called ESSAI D'UNE DISTRIBUTION GÉNÉALOGIQUE DES SCIENCES ET DES ARTS PRINCIPAUX de Chrétien Frederic Guillaume Roth (1769), which we have put up under Zoomify. The static image is a 10 MB jpeg file, available here. Browsers beware. I like this image so much, I purchased a large reproduction and had it nicely framed. Yes, the framing cost more than the reproduction. Isn't that always the case? Manuel Lima mentions the Essai to his stunning array of visualizations at Visual Complexity, which is well worth the visit, and linked it to a modern interactive representation of the Système Figuré des Connaissances Humaines by Christophe Tricot. The Encyclopédie Collaborative Translation Project has released an English translation of the Système Figuré.

Marti Hearst, Search User Interfaces

I have been reading Marti Hearst's excellent Search User Interfaces, which is fully available at http://www.searchuserinterfaces.com/. Of particular interest to me is her chapter on Information Visualization for Text Analysis. She writes "the categorical nature of text, and its very high dimensionality, make it very challenging to display graphically" and goes on to present a number of ways to handle display of text analysis results from concordances to directed graphs. This is certain something to consider for any future renovation of PhiloLogic and our related systems. We do have collocation clouds and I did a quick implementation of word frequency histograms (link) in PhiloLogic. But these are very rudimentary. Some of the examples in Hearst's a quite remarkable and we might want to model extensions of PhiloLogic on some of these.

One final note for you scribblers out here. She has a couple of entries on http://www.searchuserinterfaces.com/blog/ about how she talked her publisher (Cambridge) to let her put the book online for free and why. :-)

An important and visually compelling site/book.

Monday, May 18, 2009

Yoga for cyclists

Riding season has started again, so this should be obvious
[YouTube].

Wednesday, July 2, 2008

Shingles and Near Duplicate Detection

Sergei Vassilvitskii of Yahoo! has a useful ppt describing work to identify duplicate and near duplicate pages on the Web using shingles. Claims that 25%-40% of all WWW documents are duplicates or near duplicates. Hashing of documents cannot identify near duplicates while edit distance will not scale. Uses a hash of a small number of shingles (ngrams), calculating similarity by rate at which mini-hashes agree. Also has a useful discussion of Jaccard similarities. Talk is based on Andrei Broder's (AltaVista and Yahoo!) work, described in Identifying and filtering near-duplicate documents and previous papers cited there. There are other commercial applications of this approach, such as Equivio's near duplication identification service which uses a related similarity measure.

While I am at it, have a look at Detecting Near Duplicates in Big Data for pointers to recent work at Google on the same problem. Also, the recent International Workshop on Plagiarism Analysis, Authorship Identification, and Near-Duplicate Detection (PAN).

Tuesday, June 24, 2008

Datawocky: More data and human evaluation

Anand Rajaraman in Datawocky makes the case that more data usually beats better algorithms by reference to the NetFlix challenge and provides a little more detail in part two of the same post. He also notes that Google continues to use human evaluation as part of their search algorithm tuning in Are Machine-Learned Models Prone to Catastrophic Errors? suggesting that machine learning, based on seen instances, can suffer from the "Black Swan" problem. Finally, he makes the case, based on another blog entry, that one should Change the algorithm, not the dataset if your approach can't handle the scale of data you are throwing at it. Interesting comments all. A blog to watch.