Sunday, April 05, 2009

Unfortunate Names for 2009

I have to add Babelease Limited in the UK to my unofficial list of unfortunate names. Babel-ease. Yeah, that makes sense. I completely mis-parsed the syllables the first time...

Monday, March 23, 2009

With Apologies to Emerson

Rich are the Web-gods: who gives gifts but they?
They grope the Web for PURLs, but more than PURLs:
They pluck Force thence and give it to the wise.

Thursday, March 05, 2009

O'Reilly Media Joins the Semantic Web

O'Reilly Media (http://oreilly.com/), the current name for the geek publishing giant founded by Tim O'Reilly, has finally joined the Semantic Web.  O'Reilly's coining of the term "Web 2.0" and early misunderstandings of the Semantic Web stack lead some to think that he didn't see much value in machine readable information.  That seems to have changed, at least in within O'Reilly Labs.

O'Reilly Labs launched a Beta product last month called the O'Reilly Product Metadata Interface (OPMI), which is available at http://labs.oreilly.com/opmi.html.  The OPMI is a technical platform for the exchange of metadata between publishing trading partners.  Now that it is in RDF and publicly accessible, the rest of us can play with it, too.

It is easy to retrieve RDF/XML describing any book that O'Reilly publishes. You simply perform an HTTP GET on a URL constructed with the book's International Standard Book Number (ISBN). Every edition of a published book has an ISBN and they come in two flavors, the older 10-digit variety and the newer 13-digit version. All ISBNs issued after 1 January 2007 have been 13 digits. Some books are assigned both forms by their publishers for convenience during the transition.

For example, let's get the metadata description of an O'Reilly book I wrote, Programming Internet Email. The 13-digit ISBN for the second edition of the paperback is 9781565924796, and the 10-digit equivalent is 1-56592-479-7. The OPMI nicely works with either one, but the returned RDF uses the modern 13-digit one as canonical, as it should.

The URL for any O'Reilly book is http://opmi.labs.oreilly.com/product/ followed by its ISBN, in this case 9781565924796. The full URL is thus http://opmi.labs.oreilly.com/product/9781565924796.

An HTTP GET may be done with any Web browser, of course, or on a command line by use of the curl utility:

$ curl http://opmi.labs.oreilly.com/product/9781565924796


The returned RDF includes a wealth of information about the book. The OPMI uses four vocabulary descriptions in its RDF: Dublin Core for describing books (title, subject, language, etc), Friend-of-a-Friend (FOAF) for describing people associated with those books, the library community's MARC (MAchine Readable Cataloging) relator codes for relating people and books and the Metadata Object Description Schema (MODS) for specifying the edition of a book. MARC and MODS come from the Library of Congress and are traditionally used in library cataloging systems.

Since this metadata is on the Web, we can use standard Semantic Web query tools to query it. Using SPARQLer, a SPARQL query language processor available freely on the Web, we can query the RDF to extract bits we want. A bit of playing around makes it easy to get the author's name and the unique URI assigned to the author by O'Reilly:

prefix dc:
prefix foaf:
prefix rdf:
SELECT ?work ?authorURI ?author
FROM
WHERE {
?work dc:creator ?authorType .
?authorType rdf:_1 ?authorURI .
?authorURI foaf:name ?author
}


The results look like this:
work authorURI author
<urn:x-domain:oreilly.com: product:9781565924796.IP> <urn:x-domain:oreilly.com: agent:pdb:2495> "David Wood" @en
<urn:x-domain:oreilly.com: product:9781565924796.BOOK> <urn:x-domain:oreilly.com: agent:pdb:2495> "David Wood" @en


There are two results because the first (.IP) is the overall URI for the work in all of its possible formats. The second (.BOOK) is the book edition of the work. If this book had been published on Safari, O'Reilly's electronic publishing forum, it would also have a URL ending in ".SAF". E-books get an ".EBOOK" and Apple iPhone applications get a ".APP".

O'Reilly claims published metadata for over 1100 books, which is a pretty reasonable addition to the Semantic Web, even in Beta. Naturally, I now want O'Reilly to publish machine-readable metadata on their human-readable Web pages using RDFa. There has been no sign of that yet, though.

This content was cross-posted to Semantic Universe.

Monday, March 02, 2009

PURL Legacy Loader Now Open Source

A legacy loader is available to take old OCLC version 1 Persistent URL (PURL) database dumps and upload PURLs into the new project’s RESTful API. This is not production code, but is provided in the hope that it may be useful to operators of old PURL servers wishing to migrate to a more modern PURL server. The legacy loader has been released under an Apache 2.0 license.

To get the legacy loader, use Subversion to check it out like this:

svn co http://purlz.zepheira.com/svn/purlz/purlsbulkloader

Check out the code and follow the directions in the file README.txt.

This information is also available at the PURL Project's Download Area.

Persistent URL (PURL) Server version 1.4 Released

The PURLZ Persistent URL Server version 1.4 is now available. See the PURLZ Downloads area to get your copy now. This release improves handling of URLs with query strings and special characters. It is recommended for immediate use by all PURL server operators.

PURLs are Web addresses or Uniform Resource Locators (URLs) that act as permanent identifiers in the face of a dynamic and changing Web infrastructure. This capability provides continuity of references to network resources that may migrate from machine to machine for business, social or technical reasons. Details are available on the PURLZ community site.

Please see also the README and Release Notes for version 1.4.

Saturday, February 14, 2009

Fun with Blimps

Aidan, Mikayla and I had a blast today by attaching a digital camera to a helium blimp and flying it around our neighborhood :)

Friday, February 13, 2009

No Darwin in the South

I know I live well South of the Mason-Dixon Line. The slower pace of life here, the older attitudes and the more formal politeness is often pleasant. Sure, there are prejudices and many of the public schools aren't very good (others are, naturally). There is a lot of societal stress due to Northern migration. Virginia was even a blue state in the last election. All in all, many people from many places live in Virginia and call it home.

That's why I was shocked that my kids' school didn't even mention the 200th birthday of Charles Darwin yesterday. Neither my fifth or second grader knew who he was, or why he was famous. They know now, though. We talked about the The Voyage of the Beagle, On the Origin of Species and The Decent of Man all through dinner. Tomorrow, I plan to describe his work on worms. Kids love that sort of thing, even more than discussions of sexual selection vs. natural selection.

Shame on Fredericksburg Academy! They call themselves a college prep school? They don't even teach sex education until seventh grade! By that time, the kids have had the chance to figure it out for themselves, often in inappropriate ways. I had "the talk" with my fifth grader earlier this year. He is better for it, too. I'm honestly looking forward to talking to the head of the Lower School when she gets the rant I just sent her on Monday morning.

Monday, February 09, 2009

Desperately Seeking SKOS Vendors

A Fortune 500 customer of Zepheira's has a problem that could readily be solved with SKOS. You might think that would be sufficient to attract the attention of some tools vendors, especially since SKOS is in "last call" at the W3C and is likely to become a standard later this year. If that is so, I've missed it.

Can anyone tell me where to get decent tool support for SKOS?

Mulgara has some cool support for SKOS, as I mentioned here. Unfortunately, the state of that support still requires some care and feeding by an expert.

I approached Revelytix, hoping that they would agree to provide SKOS support in Knoodl, but they demurred until at least later this year. It should be easy for them given their existing support for OWL and their use of Mulgara.

Another alternative may be ThManager, an Open Source SKOS editor/visualizer.

Until tools vendors support SKOS directly, we are limited to existing taxonomy creation and maintenance tools, such as BiblioTech or Synaptica to build ANSI/NISO standard thesauri (Z39.19) then convert them to SKOS. For the moment, though, conversion tools seem to be in the same boat as editors.

SKOS in Mulgara's RLog

I have long been impressed by Paul's technical prowess. His recent implementation of SKOS definitions in Mulgara's RLog has done it again.

RLog is a logic programming language like Prolog that Paul created. RLog natively understands URIs and RDF's notions of subject-predicate-object relations. RLog's implementation of SKOS requires a mere 7 rules (!) once the 95 axioms are laid down. Naturally, those axioms and rules include huge chunks of RDFS and OWL.

RLog makes it easy (if you are a logic programmer) to make rules files for Mulgara's Krule rule engine. Support for RDFS has been provided in Krule for some time.

Paul has been talking about integrating RLog into Mulgara for over two years. I hope he can make that happen during 2009. Scalable or not, it is insanely cool. Until an integration happens, RLog must be run as a separate tool, as does Krule.

Friday, February 06, 2009

Ph.D. Thesis Published

My Ph.D. thesis, entitled Metadata Foundations for the Life Cycle Management of Software Systems has been published on UQ eSpace, The University of Queensland's institutional digital repository. Get your copies now while they're hot :)

Interestingly, at least to me, is that UQ eSpace is built on Fedora Commons, and therefore uses Mulgara. Sweet!

Tuesday, February 03, 2009

IET Software Journal Article Finally Published

The British journal IET Software finally published an article I wrote nearly three years ago. It was apparently published last August but I just recently found out.

The article is Towards a software maintenance methodology using Semantic Web techniques and paradigmatic documentation modelling.

The citation is:

Hyland-Wood, D., Carrington, D. and Kaplan, S. (2008, August). Towards a software maintenance methodology using Semantic Web techniques and paradigmatic documentation modelling, IET Software, 2/4, pp. 337-347

Wednesday, January 21, 2009

What is an Oracle-Mulgara Instance?

Paul pointed out a US government contract solicitation involving Mulgara. It mentions something intriguingly called an "Oracle-Mulgara instance". I am intensely curious what that is!

Persistent URL (PURL) Server version 1.3 Released

The PURLZ Persistent URL Server version 1.3 is now available. See the PURLZ Downloads area to get your copy now. This release contains substantial improvements for speed of indexing, stability and numerous bug fixes. It is recommended for immediate use by all PURL server operators.

PURLs are Web addresses or Uniform Resource Locators (URLs) that act as permanent identifiers in the face of a dynamic and changing Web infrastructure. This capability provides continuity of references to network resources that may migrate from machine to machine for business, social or technical reasons. Details are available on the PURLZ community site.

Please see also the README and Release Notes for version 1.3.

Monday, January 19, 2009

Meanwhile, Back in the Real World...

Aidan is hooked on the Mac OS X port of Nethack. Ya gotta laugh.

The Content of their Characters

Today is Martin Luther King, Jr. Day in the United States and rightfully so. We watched his "I have a dream" speech in its entirety at lunch today and I realized, in explaining his legacy to my children, how many modern-day prophets have paid the ultimate price. King, his mentor of non-violence Mohandas Gandhi and Abraham Lincoln, the three men arrayed in spirit at King's speech, were all removed from this Earth by assassins' bullets. All of them were killed for having the courage to say to small minds that people should be free.

Raised on the ideals of the American union, I was a child of King in a literal sense. King spoke at the Lincoln Memorial on the day that I was born. I grew up in prejudiced times but in hearing the conversation that he started, learned to tolerate, then to embrace, cultural differences. There are no racial differences, of course, and have not been since Neanderthals walked Europe alongside Homo Sapiens Sapiens. Such minor differences as skin color are trivial and recent evolutionary adaptations to environmental conditions that we have long worked around with forms of transportation. Culture, not race, is all that separates us.

Culture is fungible. We can change it. We have the ability if we only have the will. Do we want to live together on this increasingly tiny planet, or do we wish to let our subtle differences rip us apart? The time has come to choose. We have to work together to address the problems of our time. Climate change, energy production, medical ethics, poverty and war won't go away unless we will them to. The only way to address any of them is to live together, in peace if not always in harmony. THE challenge of our time is thus laid bare.

Tomorrow Barack Obama will become the 44th president of the United States. I am pleased that so many feel a sense of pride and accomplishment in the victory of his genes, but hope that they will remember that his victory is not about his genes, his past, or his parents. It is about our future. I, for one, support him not because he is African American, but because I believe him to be the best man for the very difficult job. I attempted to judge him, simply, not on the color of his skin, but on the content of his character.

Obama is following a dangerous path. He will need to ignore his own rock star status, to avoid offers from young women, to avoid the corrupting influences of Washington. He will need to avoid assassins' bullets. If he lives, if he stays sane, if he can just do what he has set out to do, he just might become truly great. I hope he can. I hope we can follow.

Wednesday, November 19, 2008

Mulgara on the Cloud

Chris Wilper reportedly put a Mulgara instance on Amazon EC2 and loaded a quarter-billion triples into it, just for fun. The data he used was generated from Paul Gearon's numbers RDF generator. The generator creates facts about numbers and keeps creating RDF until you stop it.

Chris was using the new XA version 1.1 storage layer for Mulgara. The quarter-billion triples loaded in about a day, which is about the same loading performance that Paul has seen on his laptop. The size of the indexes were about 4.6 gigs compressed and 51 gigs uncompressed. Paul notes that the XA version 2 storage layer will store strings and URIs much more efficiently and thus reduce the index sizes considerably.

Way to go, guys! We can't wait for XA2!

Tuesday, November 18, 2008

New PURL Code in the Wild

Several PURL installations have been seen on the 'net using the new code base:
  1. NeuroCommons
  2. The National Center for Biomedical Ontology (NCBO)
  3. Semantic Report
  4. Zepheira (naturally)
Others include at least one startup that doesn't have a Web site yet (GRACE Research Corporation) and a couple of others still in stealth. Not bad for a project that never had an official launch. Keep 'em coming, folks. I have some hope that OCLC will join that list soon, as well as a couple more startup companies.

PURL Server v1.2 Released

I released the new Persistent URL (PURL) server, version 1.2 over the weekend. This release fixes a major bug in version 1.1 whereby PURLs would not resolve for those not having a cookie set. Minor upgrades include better error reporting and URL handling.

All users of earlier PURL servers are encouraged to upgrade immediately.

Binary (JAR) and source code downloads are available from purlz.org.

Saturday, November 08, 2008

The Perfect 11-year-old Boy Birthday Party

4:00 PM: Archery (supervised, of course)

4:45 PM: A few minutes free time on a trampoline

5:00 PM: Fencing lesson (even more stringently supervised, with cardboard boxes as opponents)

6:15 PM: Make your own pizza for dinner

7:15 PM: Communal game of Crossfire on a map created by the birthday boy himself

8:30 PM: Cookies

8:35 PM: Pass-the-parcel (an Aussie game for our Aussie son)

8:40 PM: Open presents

8:50 PM: Movie night (The Princess Bride)

10:30 PM: Sleep over

8:00 AM: Breakfast of chocolate-chip pancakes

9:00 AM: Pickup by parents

Wednesday, November 05, 2008

My Last Degree

Tonight, I passed the oral defense for my Ph.D., having submitted my thesis in July. It's all administration from here :)