Showing posts with label PURL. Show all posts
Showing posts with label PURL. Show all posts

Tuesday, April 07, 2009

OCLC PURL Server Migrates to PURLZ

The Online Computer Library Center (OCLC) migrated http://purl.org to the PURLZ software this morning at 6 AM US EST (GMT -5).

I'll admit to some frustration that the legacy PURLs were not tested completely and some errors remain. We are working with them to iron out the relatively few remaining problems with the legacy data migration. Most of the legacy data seems to be working as expected.

Update: Nope. They rolled back again. Sigh. Maybe next time they will do it right.

Monday, March 02, 2009

Persistent URL (PURL) Server version 1.4 Released

The PURLZ Persistent URL Server version 1.4 is now available. See the PURLZ Downloads area to get your copy now. This release improves handling of URLs with query strings and special characters. It is recommended for immediate use by all PURL server operators.

PURLs are Web addresses or Uniform Resource Locators (URLs) that act as permanent identifiers in the face of a dynamic and changing Web infrastructure. This capability provides continuity of references to network resources that may migrate from machine to machine for business, social or technical reasons. Details are available on the PURLZ community site.

Please see also the README and Release Notes for version 1.4.

Wednesday, January 21, 2009

Persistent URL (PURL) Server version 1.3 Released

The PURLZ Persistent URL Server version 1.3 is now available. See the PURLZ Downloads area to get your copy now. This release contains substantial improvements for speed of indexing, stability and numerous bug fixes. It is recommended for immediate use by all PURL server operators.

PURLs are Web addresses or Uniform Resource Locators (URLs) that act as permanent identifiers in the face of a dynamic and changing Web infrastructure. This capability provides continuity of references to network resources that may migrate from machine to machine for business, social or technical reasons. Details are available on the PURLZ community site.

Please see also the README and Release Notes for version 1.3.

Tuesday, November 18, 2008

PURL Server v1.2 Released

I released the new Persistent URL (PURL) server, version 1.2 over the weekend. This release fixes a major bug in version 1.1 whereby PURLs would not resolve for those not having a cookie set. Minor upgrades include better error reporting and URL handling.

All users of earlier PURL servers are encouraged to upgrade immediately.

Binary (JAR) and source code downloads are available from purlz.org.

Monday, February 18, 2008

Beyond Redirection: Rich and Active PURLs

It has been a while since I posted about the new Persistent URL work being done by Zepheira and OCLC. That is partially because the work has gone much more slowly than we had planned. We have been actively gold plating the new PURL server because we are trying to satisfy several communities. The result, though, is laying the foundations for new services for the Web.

The key to the new PURL service is the typing of PURLs. The existing public PURL service at OCLC returns an HTTP 302 (Found) response, causing a Web client to redirect to another URL. The new PURL server allows PURLs to return one of several status codes (301, 302, 303, 310, 404 and 410).

We can go even further, though. The original PURL server had some internal concepts like "cloning" a PURL (basing a new PURL definition on an existing one) and "chaining" a PURL (allowing multi-person management of a URL resolution process). Combining these concepts with the choice of HTTP response codes got us thinking about arbitrary types of PURLs.

We have been experimenting with types of PURLs that combine with other services. A "Rich" PURL, for example, is the combination of a PURL and metadata. We have a prototype service that combine strong identifiers with rich metadata, providing the building blocks for other semantic applications. Rich PURLs are a combination of two related services: A PURL server for management of the resolution and RDF (or other metadata format) file hosting.

The W3C TAG finding regarding the use of HTTP 303 responses seems to suggest that we could use rich PURLs in interesting ways. For example, we could do the following:

A 301 PURL pointing to an RDF resource == Metadata about an information resource
A 303 PURL pointing to an RDF resource == Metadata about a non-information resource (i.e. a physical or conceptual resource)

That usage would be consistent with the TAG finding, even if it goes a bit beyond it.

NB: You can tell if you get a PURL by looking at the PURL header in an HTTP response.

But wait, there's more. What if a PURL pointed to a Web service (in the sense of dynamic content, not necessarily limited to SOA Web Services, but including them)? The combination of a Rich PURL and a metadata reference to a Web service yields an "Active PURL". That is, an Active PURL is a PURL naming a graph of metadata describing a Web service.

Consider a simple Web service like an RSS feed. Placing an Active PURL in front of that feed allows you to describe how that feed should be handled. You could name the facets that you want to use to make an Exhibit or provide any other presentation advice you desired.

Alternatively, an Active PURL might itself by a sort of Web service that provides dynamic metadata about another Web service and can either serve the metadata or redirect to its target service. Such an Active PURL could be used to name a SPARQL graph, accept query parameters for it and return metadata about it, such as a count of results or information on the meanings of columns in the result set. I believe that named graphs are very handy things and something that we as a community are paying inadequate attention to. Given that SPARQL may be the query language that finally integrates our silos of relational databases, fronting them with Active PURLs seems like a promising line of research.

Lists of URLs as proposed by Stu Weibel would be easy to implement as an Active PURL.

Perhaps the most interesting use of Active PURLs to enterprises might be the ability to provide standardized RDF metadata about SOA Web Services as well as relational databases. UDDI is so broken, we might as well fix it with existing SemWeb standards. That is not a new idea, but the application of Active PURLs to the problem is, I think.

Wednesday, July 11, 2007

Press Release for PURL Code Rewrite

Zepheira, OCLC and the W3C announced today a project that we have been working on for a month; the complete re-write of the software behind the Persistent URL (PURL) service.

The full press release is here.

I am particularly happy to say that we are working toward a proper Open Source (Apache 2) release of the new code, and enhancements to deal with important issues for the Semantic Web (such as support for HTTP Range-14 return codes). OCLC has also asked us to help build a community around the code to assist both its maintenance and its future direction. That is a refreshing change and should be welcomed by all.

I can't wait to start building services on top of this stuff.

Monday, June 04, 2007

Lots of Zepheira Press

Zepheira has been getting a lot of press recently. The Zepheira news page certainly has links to more podcasts and respectible publications than any other company I have have the privilege to be part of.

Investors Business Daily came out today with an interview of several key SemWeb leaders, including Eric Miller. The article itself expires after today (!) so I have created a PURL for it so we can eventually move the redirection when IBD pulls their head out.

The article was typical journalistic fare, attempting to be "balanced" in the same way that Fox News is. That is done by showing both sides of a story - even when there are not, in fact, two sides. We see this sort of nonsense with science reporting when reporters go increasingly out of their way to find a few scientists willing to publicly doubt the phenomenon of climate change or willing to say that nuclear fusion will happen commercially within the decade. In the case of the Semantic Web, the same old doubters of 1999 come out to play and get quoted every time, ignoring the clear commercial successes of the last few years.

It is interesting to see Tim O'Reilly routinely quoted in the press now putting the SemWeb in a more positive light than last year: "The Semantic Web is the idea of marking up computer information in such a way that computers can infer meaning from it."

The journalist managed to get basic facts wrong, too, as when he says, "Internet protocols make it hard to search for many basic items on a Web site, such as a simple address or phone number." or "A resource description framework (RDF) and Web ontology language (OWL) are new technologies that can solve that problem. They serve as a kind of wrapper or tag to describe the data inside." RDF and OWL are hardly new (both were standardized in 2004 and pretty stable years before) and of course the Internet protocols have nothing to do with linguistic searching.

He also gets some things right: "This new vocabulary lets computers find and access data on their own. The goal is letting the machines perform rote tasks to gather information and merge the results.", and he mentions products from Oracle and Adobe.

Eric is quoted reasonably. That is a relief. "These Web standards should help companies spot new relationships among huge sets of data and use the findings for better conclusions about their business, says Eric Miller, president of Web startup Zepheira."

I rather liked the MySpace example and wonder if it came from Eric:


For instance, MySpace might let personal pages share information with the pages of relevant friends or colleagues in the social network.

Take someone whose MySpace page describes a fondness for vintage jazz. By entering that information once, that person could automatically be linked to others who share the same interest.

Furthermore, that information could be applied to future Web searches for new music releases. In effect, using metadata could become a way to make MySpace "truly mine," said Miller.

"This means there is a much more flexible, personalized integration point to really connect people," he said. "The notion here is to enter data just once, but to use it often."