275: URIs, IRIs, and basic paths

by Duncan Guthrie

Status

This SRFI is currently in draft status. Here is an explanation of each status that a SRFI can hold. To provide input on this SRFI, please send email to srfi-275@nospamsrfi.schemers.org. To subscribe to the list, follow these instructions. You can access previous messages via the mailing list archive.

Abstract

This SRFI proposes a programming interface for working with RFC 3986 universal resource identifiers (URIs), as well as RFC 3987’s generalisation to internationalised resource identifiers (IRIs). This document defines record types, normalization procedures, and conversion between URIs and IRIs. The proposal also specifies a basic programming interface for working with paths in isolation, as this a pervasive usage of URIs and IRIs. Finally, we contribute a test suite to specify the behaviour of normalization with respect to relative references, which has in the past been a source of divergence between implementations.

Issues

none so far

Table of contents

Rationale

RFC 3986 [1] describes an abstract syntax for uniform resource identifiers (URIs) as well as their relative references, which allow documents to be authored without knowing the final publishing location. RFC 3987 [2] defines IRIs, which generalise URIs to Unicode by defining an interpretation of percent-encoded bytes as sequences of UTF-8 octets.

URIs and IRIs are widely used to denote resources across the world wide web, and are the basis of a number of web standards. It is critical to present a programming interface for manipulation of different components in isolation (e.g. paths, hostnames), as working with an URI’s coarse string representation directly is error-prone. RFC 3986 defines an abstract syntax and hence conveniently forms the basis of such a programming interface. Further, RFC 3987’s generalisation to internationalised identifiers is a natural extension, allowing us to faithfully denote resources in a number of languages, using the universal character set. Indeed, IRIs form the basis of modern standards like Resource Description Framework (RDF), a widespread formal model for metadata interchange and knowledge representation. We hence require support for both URIs and IRIs.

Finally, paths are one of the most prevalent applications of URIs and IRIs. We develop paths as an object disjoint from other Scheme types with distinct segment structure, rather than adopting a string representation, because normalisation of paths is a key source of divergence in URI and IRI implementations to date. This object is located in a basic path library. This library structure also lets us expose procedures such as normalisation and merging of base and relative paths directly to the programmer.

Relative references and polymorphic procedures

RFC 3986 and 3987 distinguish URIs and IRIs, respectively, from their relative references, which must be resolved against a base URI or IRI in order to be used. The main application of relative references is to allow one to refer to resources and author documents without knowing the final publishing location. For example, a graph database may produce RDF documents specified in RDF/XML, where resources are denoted with IRI-relative references, assuming that these documents would be interchanged with another system which mints full IRIs with respect to its hosting location.

Somewhat confusingly, RFC 3986 defines "URI-reference" as the most common usage of URI: either an URI or a relative reference. We follow this common usage by developing polymorphic getters and setters which work on URIs and relative references, with the predicate uri? holding true for both URIs and relative references. This procedure would hence correspond to testing for an RFC 3986 "URI-reference". (RFC 3987 follows an identical convention for IRIs and their relative references, so the same approach is followed for IRIs.)

The most divergent behaviour between widespread implementations of URIs and IRIs has been with respect to normalization of relative references. We explicitly avoid defining path segment normalization for relative references (regardless of whether the path is absolute or relative) because the behaviour is largely undefined by RFC 3986 and 3987 or any current RFC. See the section on normalization for more details and the test suite.

A record type presentation

URIs and IRIs are structured data. Of course, record types are convenient as they generate dedicated getters and setters. More importantly, however, we think that record structures are necessary in this case for normalization procedures to be structure-preserving. Specifically, normalization procedures may alter the path, such that its segments may be mistaken for other parts of the URI or IRI, such as the <://> portion separating scheme from authority, or for the hostname. Updates to authority components or to the path need to similarly ensure that they produce valid URIs and IRIs when serialised.

We argue that a string representation makes it too easy to inadvertently modify the structure, because normalization of individual components may yield a string which, when parsed again, is interpreted differently with respect to the structure. A good example of this is the restrictions on paths given an authority, because a non-empty authority is denoted in an URI using two slashes, which are characters which may also appear in paths.

A string representation of an URI or IRI (procedures uri->string and iri->string) is produced by concatenating string representations of the individual fields (scheme, hostname &c.), with the expected separators between components. For efficiency, no assumption should be made that the contents of individual fields can be checked at this point, which typically would involve additional, redundant parsing.

Pure and impure interfaces

We require implementations to provide a purely-functional interface to URIs and IRIs, but not an in-place interface. The reasoning for this is that if implementors choose data structures optimised to purely-functional programming, it is more cumbersome to create an impure interface, whereas it is not as cumbersome for implementors to create a (potentially inefficient) purely-functional interface by copying the URI or IRI before updating in-place.

If an implementation provides disjoint mutable and immutable URI and IRI variants, then it is an error to call the in-place setters on an immutable variant, but the in-place variants should otherwise have equivalent error handling to the purely-functional variants.

Error behaviour of setters

Our design is to provide getters and setters which abstract away the internal representation, working on string representations of URI components. More generally, we suspect that existing implementations largely omit setters because preserving internal consistency is fairly cumbersome on the implementor and programmer, with validity of authority components being defined with respect to the path, and vice versa. The specific challenge is to ensure that setting a given component would not violate the URI grammar, as this may elicit a flat string representation which would have a different structure when parsed again.

In order to set one of the three authority components safely on an IRI or URI, then the following conditions must hold with respect to the IRI or URI’s existing path shape:

  1. If the authority component becomes set (at least one of the three components are not #f), rfc-path-abempty? must hold against the path.
  2. If the authority component becomes unset (all of the three components are #f), and the IRI or URI is non-relative, then either rfc-path-absolute?, rfc-path-rootless? or rfc-path-empty? must hold against the path.
  3. If the authority component becomes unset (all of the three components are #f), and the IRI or URI is a relative reference, then either rfc-path-absolute?, rfc-path-noscheme? or rfc-path-empty? must hold against the existing path component.

Similarly, in order to set the path component safely, then the following conditions must hold with respect to the three authority components:

  1. If either of the authority components above are set, then rfc-path-abempty? must hold against the path to be set.
  2. If no authority component is set, and the IRI or URI is non-relative, then either rfc-path-absolute?, rfc-path-rootless? or rfc-path-empty? must hold against the path to be set.
  3. If no authority component is set, and the IRI or URI is a relative reference, then either rfc-path-absolute?, rfc-path-noscheme? or rfc-path-empty? must hold against the path to be set.

Helper predicates for checking whether setting these components are located in the utility library: user-settable?, host-settable?, port-settable? and path-settable?.

Normalization

We support three, scheme-independent normalization procedures:

  1. Case normalization, in which U.S. ASCII may be case-insensitive, and in which escaped octets have a canonical upper-case form.
  2. Escape normalization, in which percent-encoded bytes may be decoded, or characters may be encoded as escapes.
  3. Path segment normalization, in which dotted segments (‘.’ and ‘..’) are eliminated.

Case normalization

The scheme and host components of both URIs and IRIs are considered case-insensitive, with other components considered case-sensitive. For URIs, the repertoire of characters is within U.S. ASCII., whereas for IRIs, Unicode characters may appear. Nonetheless, for both URIs and IRIs, only U.S. ASCII is case-insensitive, with characters like "É" never being normalized to a lower-case form like "é".

Escapes take the form of percent-encoded bytes, appearing as % 0-F 0-F (hexadecimal digits) in URIs and IRIs. These hexadecimal digits have a canonical upper-case form. For example, %cf should be normalized to %CF. In practice, the sample implementation always parses these into a canonical form as it represents these internally as octets, not in the original string form. If implementations do preserve the original form, they must always support normalization into the canonical upper-case form.

Escape (percent-encoded byte) normalization

The interpretation of escapes (percent-encoded bytes) differs for URIs and IRIs.

For URIs, these octets may individually be decoded to ASCII characters. Essentially, this occurs if a character is not in the URI reserved range, and if it is permissible within a given URI component (e.g. path). Characters in the reserved range, if encountered in the clear, must never be escaped. Conversely, characters not in the reserved range, and which are not permissible within a given URI component, are encoded as a series of percent-encoded bytes corresponding to the character's UTF-8 octets. This bears particular mention because, while other character encodings are valid URIs, the RFC 3986 specification specifically requires this encoding, which enables compatibility with the closely related RFC 3987 specification for IRIs.

IRIs not only generalise the range of characters permissible within the IRI to certain Unicode ranges, but also interpret sequences of percent-encoded bytes as octets in UTF-8, which may be decoded into legal characters in the universal character set. Additionally, because all URIs are valid IRIs, the normalization of IRIs with respect to percent-encoded bytes is essentially the same as conversion from an URI to an IRI.

Path segment normalization

Path segment normalization is structurally identical for both URIs and IRIs. It interprets an entire path with respect to two control sequences: ‘.’ (current working directory) and ‘..’ (upper working directory), similar to UNIX paths. Unlike the other two normalization procedures, path segment normalization is undefined for relative references, because the path portion is not meaningful except during relative reference resolution.

It should be noted that path segment normalization is defined for non-relative IRIs with relative paths, such as a URN like <foo:a/b/../.././../../e>. This is a major source of diversion between RFC 3986 implementations. This appears to arise from reusing the Remove Dot Segments procedure as defined in RFC 3986 verbatim. In relative reference resolution, this procedure is never called on relative paths, only on absolute paths.

For example, for the (non-relative) IRI <foo:a/b/../.././../../e>, implementations using the RFC 3986 procedure verbatim get <foo:/e>, whereas other implementations get <foo:e>. In the former camp are implementations like the Erlang/OTP system’s built-in uri_string [3], and Guile-RDF [4], whereas in the latter camp are implementations like Haskell’s network-uri [5] and Chicken’s uri-generic [6] and uri-common [7] and libraries. We are in the latter camp, and a fixed Remove Dot Segments procedure might be implemented as follows. The sample implementation internally represents raw segments as a vector of SRFI 160 [10] u16vector segments, and uses the SRFI 262 pattern matcher [11].

(define (remove-dot-segments path)
  (let* ([undotted
          (remove-dot-segments/list (path-raw-segments path))]
         [segments
          (list->vector undotted)])
    (cond [(absolute-path? path)
           (make-raw-absolute-path segments)] ;; implementation-specific
          [(and (relative-path? path)
                (fx=? 1 (vector-length segments))
                (u16vector-empty? (vector-ref segments 0)))
           ;; relative path with single empty segment is not meaningful:
           (make-raw-relative-path (vector))] ;; implementation-specific
          [else
           (make-raw-relative-path segments)]))) ;; ''

(define (current-wd? seq) (u16vector= seq (u16vector #x2E)))

(define (parent-wd? seq) (u16vector= seq (u16vector #x2E #x2E)))

(define (remove-dot-segments/list segments)
  (let loop ([ps (vector->list segments)]
             [trailing-slash? #f]
             [lst '()])
    (match ps
      ['()
       (if trailing-slash?
           (reverse (cons (u16vector) lst))
           (reverse lst))]
      [(cons (? current-wd?) rst)
       (loop rst #t lst)]
      [(cons (? parent-wd?) rst)
       (loop rst #t (if (pair? lst) (cdr lst) lst))]
      [(cons x rst)
       (loop rst #f (cons x lst))])))

Additionally, the above suggested implementation of remove-dot-segments is considerably clearer than the stack-based description in RFC 3986, with fewer pattern-matching clauses required.

Relative reference resolution and relativization

Relative reference resolution resolves an URI or IRI against a (non-relative) base URI or IRI. This SRFI proposal simply names these procedures resolve-iri and resolve-uri as the "relative reference resolution" language in RFC 3986 and 3987 specifically refers to the URI-reference and IRI-reference (see above). Additionally, this naming is consistent with existing libraries such as Erlang's uri_string module [3] and Java's URI module [12].

The inverse of this procedure takes a non-relative URI or IRI and a base URI or IRI, and produces a relative reference which, given the base, would produce that non-relative URI or IRI. This is not defined in RFC 3986 or 3987, but is fairly widespread, such as relative-from found in Haskell's network-uri [5] and Chicken's uri-generic [6], and relativize found in Java's URI module [12]. The Haskell and Chicken libraries arguments are flipped for consistency with their direction, whereas Java's class methods belong to the base URI. This SRFI proposal names these procedures relativize-iri and relativize-uri, which follows Java's naming and argument convention, although the behaviour and test cases are based on the Chicken library:

Relativization proceeds as follows:

  1. If the scheme components differ, then the non-base argument is copied.
  2. If the authority components differ, then a relative reference is returned based on the non-base argument's fields.
  3. If the path components differ, then the two paths (after normalizing dotted segments) are compared in order to calculate a path relative to the base argument's path. A relative reference is returned using this relative path, with all three authority components unset (set to #f), and with the query and fragment of the non-base argument.
  4. If the query components differ, then a relative reference is returned with all three authority components unset, the empty path, and with the query and fragment of the non-base argument.
  5. Exhausting the above cases implies that the path components are equal (and that path-rfc-abempty? holds. A relative reference is returned in which the authority components are unset, query is unset, and with the following path:
    1. If there are no segments (single slash), use the path as-is
    2. If the final segment is empty, then use an absolute path with one segement, a single dot (‘.’)
    3. If the final segment is not empty, then use an absolute path with that final segment.

Calculating that relative path, on the level of segments behaves as follows, or an equivalent procedure:

  1. If the base path is empty, then copy the non-base path.
  2. If the non-base path is empty, return an empty relative path.
  3. Partition both path arguments into a sequence of non-final segments, and the final segment.
  4. Compare the two sequences of non-final segments to find the number of identical segments, k.
  5. Replace the first k segments in the non-base initial segment sequence with k  ‘..’ segments. This sequence is our candidate relativized path segments.
  6. This candidate sequence of relativized segments may be empty, and if so:
    1. If the final segment of the original two paths was equal, then return the an empty sequence of segments.
    2. If the final segment was empty, return a sequence with one segment, a single dot (‘.’)
    3. Otherwise, return a sequence with one segment, the final segment of the non-base path.
  7. Otherwise, if the candidate sequence of relativized segments was non-empty, and its last segment was dotted (‘.or..’):
    1. If the final segment of the original non-base path was empty, use the candidate sequence as-is.
    2. Otherwise, append the final segment of the original non-base path to the end of the candidate sequence.

Finally, these segments are wrapped by a path object. The initial segment may be empty, in which case, the path object is absolute with segments excluding the initial segment. Otherwise, the path is a relative path using the segments calculated above. Additionally, the relativization procedures check that the path was settable given the authority component, and will prepend ./ to paths with a leading double slash given no authority component being set (see here).

Basic paths

We additionally specify a basic path sub-library. Paths can be considered vectors of segments which are tagged with whether they have a leading slash (an absolute path). The character set allowed within these paths is equivalent to the characters permissible within an URI or within the range of the universal character set (UCS) permitted within an IRI path segment. This restriction is important because it excludes a number of control characters and characters which can never be typed.

The basic path library specifies the basic operations of path creation from sets of segments (build-path), subscripting (path-ref), functional updates (path-update), and count of segments (path-length). We go a little further and provide utility procedures for path shape (the variants in the RFC 3986 and 3987 ABNF), and for copying paths. Finally, this design allows us to explicitly expose in the programming interface generalist procedures inherited from RFC 3986, such as merge-paths and remove-dot-segments.

R6RS conditions

For each procedure, error-handling behaviour is described in terms of exceptions to be raised and their irritants, either as an assertion violation, or as an error. R6RS conditions may not be available, in which implementations may simply signal an error. However, if these are available, then implementations should raise a compound condition as follows:

With access to R6RS conditions, one might implement the encode-string procedure as follows. The first call to the R6RS assertion-violation procedure raises a compound condition of &who, &irritants, &assertion-violation and &error (as well as &message), where the who portion is the symbol encode-string, and the single irritant is the string argument.

(define encode-string
  (case-lambda
    [(str)
     (encode-string str (char-set-complement char-set:iri-unreserved))]
    [(str cset)
     (cond [(not (string? str))
            (assertion-violation 'encode-string "not a string" nonstr)]
           [(not (char-set? escape-these))
            (assertion-violation 'encode-string "not a character set" noncset)]
           [else
            ...])]))

An R7RS-small implementation might implement it instead as follows:

(define encode-string
  (case-lambda
    [(str)
     (encode-string str (char-set-complement char-set:iri-unreserved))]
    [(str cset)
     (cond [(not (string? str))
            (error "not a string" nonstr)]
           [(not (char-set? escape-these))
            (error "not a character set" noncset)]
           [else
            ...])]))

Treatment of IP literals

In RFC 3986 and 3987, IP literals (namely IPv6 addresses) appear in the host portion of an URI or IRI enclosed by square brackets. This SRFI proposal does not normalise these internally to an IP literal object or similar, with the IP literal returned by getters for the host portion enclosed by square brackets. This has the advantage that getters and setters exchange the same (bracketed) IP literal representation, receiving applications likely need to remove the brackets before usage of IP literals returned by this library's setters.

Argument restrictions and convention

This specification follows the R6RS procedure entries convention. In addition to the naming conventions specifying type restriction for arguments where they are used, we add the following:

irinon-relative IRI or IRI relative reference
non-relative-irinon-relative IRI
relative-iriIRI relative reference
urinon-relative URI or URI relative reference
non-relative-urinon-relative URI
relative-uriURI relative reference
pathrelative or absolute path object
relative-pathrelative path object
absolute-pathabsolute path object

Notation and convention

Although we support both URIs and IRIs, for brevity we primarily describe behaviour for IRIs, and omit descriptions of the equivalent procedures for URIs where they behave the same. This works because URIs and IRIs are structurally identical, with the divergence between RFC 3986 and 3987 arising from the generalisation of the character set, and the treatment of normalization.

Library references are in the form (srfi :275 <sub-library>) consistent with SRFI 97 [9].

Finally, throughout this document, in examples we enclose IRIs and URIs in angle brackets, e.g. <http://example.org>. An object is assumed to be either an IRI or URI depending on the producing procedures, e.g. string->iri.

Index of exported procedures and variables

Constructors
string->iri     string->non-relative-iri     string->rfc-iri-reference     string->rfc-absolute-iri
string->uri     string->non-relative-uri     string->rfc-uri-reference     string->rfc-absolute-uri
get-iri     get-non-relative-iri     get-rfc-iri-reference     get-rfc-absolute-iri
get-uri     get-non-relative-uri     get-rfc-uri-reference     get-rfc-absolute-uri
empty-iri    empty-uri
Predicates
iri?     relative-iri?     non-relative-iri?     rfc-iri-reference?     rfc-absolute-iri?
uri?     relative-uri?     non-relative-uri?     rfc-uri-reference?     rfc-absolute-uri?
Equivalence and conversion
iri-equal?    iri-eqv?
uri-equal?    uri-eqv?
string->iri     iri->string     string->uri     uri->string
encode-string    decode-string
iri->uri    uri->iri
Getters
iri-scheme
iri-user    iri-host    iri-port    iri-authority    iri-username+password
iri-path    iri-path-string
iri-query    iri-fragment
uri-scheme
uri-user    uri-host    uri-port    uri-authority    uri-username+password
uri-path    uri-path-string
uri-query    uri-fragment
Setters
update-iri-scheme
update-iri-user    update-iri-host    update-iri-port    update-iri-authority
update-iri-path    update-iri-query    update-iri-fragment
update-uri-scheme
update-uri-user    update-uri-host    update-uri-port    update-uri-authority
update-uri-path    update-uri-query    update-uri-fragment
Relative reference resolution
resolve-iri    relativize-iri
resolve-uri    relativize-uri
Normalisation
normalize-iri-case     normalize-iri-escape     normalize-iri-path
normalize-uri-case     normalize-uri-escape     normalize-uri-path
Path constructors
build-path     vector->relative-path     vector->absolute-path
string->path     string->relative-path     string->absolute-path
empty-relative-path     empty-absolute-path
vector->rfc-path-rootless     vector->rfc-path-noscheme     vector->rfc-path-abempty     vector->rfc-path-absolute
string->rfc-path-rootless     string->rfc-path-noscheme     string->rfc-path-abempty     string->rfc-path-absolute
Path serialisation
path->string    path-segments
Path predicates
path?     path-string?     relative-path?     absolute-path?     empty-path?
iri-path-relative?     iri-path-absolute?     iri-path-empty?    
uri-path-relative?     uri-path-absolute?     uri-path-empty?    
rfc-path-rootless?     rfc-path-noscheme?     rfc-path-abempty?     rfc-path-absolute?     rfc-path-empty?
iri-path-rfc-rootless?     iri-path-rfc-noscheme?     iri-path-rfc-abempty?     iri-path-rfc-absolute?     iri-path-rfc-empty?
uri-path-rfc-rootless?     uri-path-rfc-noscheme?     uri-path-rfc-abempty?     uri-path-rfc-absolute?     uri-path-rfc-empty?
Path operations
path-length     path-ref     path-update     path-equal?     remove-dot-segments     merge-paths
iri-path-segment-ref     iri-path-segment-update     iri-path-segments
uri-path-segment-ref     uri-path-segment-update     uri-path-segments
Utilities
utf8->string/raise     percent-encoding->string     percent-encoding->u8-list     percent-encoding->reverse-u8-list
username+password
user-settable?     host-settable?     port-settable?     path-settable?
Character sets
char-set:iri-unreserved     char-set:ucschar     char-set:iri-private     char-set:uri-unreserved
char-set:gen-delims     char-set:sub-delims    char-set:reserved
char-set:scheme
char-set:iri-userinfo    char-set:uri-userinfo
char-set:iri-reg-name    char-set:uri-reg-name
char-set:iri-segment    char-set:uri-segment
char-set:iri-query    char-set:uri-query
char-set:iri-fragment    char-set:uri-fragment
char-set:iri-allowed     char-set:uri-allowed     char-set:path-allowed

Specification

IRI programming interface

(srfi :275 iri)
procedure
(iri? obj)

Returns #t if obj is an IRI. Returns #f otherwise.

(srfi :275 iri)
procedure
(non-relative-iri? obj)

Returns #t providing that obj is a non-relative IRI (has a scheme component). Returns #f otherwise.

(srfi :275 iri)
procedure
(relative-iri? obj)

Returns #t providing that obj is an IRI relative reference (no scheme component). Returns #f otherwise.

Examples:

(define example-IRI (string->iri "/ex#IRI"))

example-IRI → </ex#IRI>

(iri? example-IRI) → #t
(non-relative-iri? example-IRI) → #f
(relative-iri? example-IRI) → #t
  
(srfi :275 iri)
procedure
(rfc-iri-reference? obj)

Exactly iri?, which already corresponds to RFC 3987’s IRI-reference production.

(srfi :275 iri)
procedure
(rfc-absolute-iri? obj)

Returns #t providing that obj is a non-relative IRI and that its fragment component is #f. This corresponds exactly to RFC 3987’s absolute-IRI production. Returns #f otherwise.

Examples:

(define example0 (string->iri "http://example.org/ex?cond"))
(define example1 (string->iri "http://example.org/ex#title"))
(define example2 (string->iri "//example.org/ex?cond"))
(define example3 (string->iri "//example.org/ex#title"))

(map rfc-absolute-iri? (list example0 example1 example2 example3))
→ (#t #f #f #f)
  
(srfi :275 iri)
procedure
(empty-iri)

Helper procedure which takes no arguments and constructs a relative reference with all components set to #f, i.e. <> or the result of parsing "".

(srfi :275 iri)
procedure
(iri-scheme non-relative-iri)

Retrieve the scheme component of non-relative-iri as a string. Unlike the procedures to follow, this procedure never returns #f as this would imply that the scheme is unset, i.e. a relative reference.

(srfi :275 iri)
procedure
(iri-user iri)
(iri-host iri)
(iri-query iri)
(iri-fragment iri)

Retrieve the respective RFC 3987 fields of iri as either a string, or #f. An empty field is distinct from an unset one, for example given an IRI with bare ?, iri-query would return "", whereas if ? had been omitted, iri-query would return #f.

(srfi :275 iri)
procedure
(iri-port iri)

Retrieve the RFC 3987 port field of iri as either a non-negative integer, or #f.

(srfi :275 iri)
procedure
(iri-path iri)

Retrieve the RFC 3987 path field of iri as a path object as described in the path sub library. Unlike the other fields, #f is never returned.

Examples for the above seven component-specific procedures:

(define example-IRI (string->iri "http://example.org:80/ex#IRI"))

example-IRI → <http://example.org:80/ex#IRI>

(iri? example-IRI) → #t
(non-relative-iri? example-IRI) → #t

(iri-scheme example-IRI) → "http"
(iri-user example-IRI) → #f
(iri-host example-IRI) → "example.org"
(iri-port example-IRI) → 80
(path-segments (iri-path example-IRI)) → #("ex")
(iri-query example-IRI) → #f
(iri-fragment example-IRI) → "IRI"
  
(srfi :275 iri)
procedure
(iri-authority iri)

Get the authority components of iri (user, host and port). This procedure returns #f if none of the three authority components are set, else all three as multiple values.

(srfi :275 iri)
procedure
(iri-username+password iri)

Helper procedure which splits the user field of iri at the first colon, producing strings corresponding to username and password as two values. If there is no colon then the whole user field is returned as first value, and #f as second. If the user field is not set, then #f and #f are returned. A colon appearing as a percent-encoded byte (%3A) is not considered a delimiter between username and password.

Examples:

(iri-username+password (string->iri"//foo:bar:qux@host"))   → "foo"        "bar:qux"
(iri-username+password (string->iri"//foo%3Abar:qux@host")) → "foo%3Abar"  "qux"
(iri-username+password (string->iri"//@host"))     → ""    #f
(iri-username+password (string->iri "//"))         → #f    #f
(iri-username+password (string->iri "//foo@host")) → "foo" #f
  

We also provide a generic helper procedure for processing strings retrieved from IRIs or URIs user fields.

(srfi :275 iri)
procedure
(iri-path-string iri)

In contrast to iri-path, return a string representation of the path of iri, which may be the empty string.

Examples:

(define example0 (string->iri "http://example.org:80/ex#IRI"))
=(define example1 (string->iri "http://example.org:80#IRI"))
(define example2 (string->iri "//a"))

(path-segments (iri-path example0)) → #("ex")
(iri-path-string example0) → "/ex"

(path-segments (iri-path example1)) → #()
(iri-path-string example1) → ""

(path-segments (iri-path example2)) → #()
(iri-path-string example2) → ""
  
(srfi :275 iri)
procedure
(update-iri-scheme non-relative-iri string)

Purely-functional setter for the scheme component of non-relative IRI non-relative-iri, encoding the string string. During parsing, an error is raised if a character cannot be contained within a scheme at that position, with the character, its position in the IRI, and the IRI as irritants. Unlike the other setters for IRI components to follow, setting the scheme to #f or the empty string is not possible becasue it would imply that the IRI is a relative reference.

(srfi :275 iri)
procedure
(update-iri-user iri obj)
(update-iri-host iri obj)
(update-iri-port iri obj)

Purely-functional setters to set a respective authority components of IRI iri to obj. These procedures check that the authority component after being set to obj is not in conflict with the shape of the existing path. For instance, relative paths cannot usually be set when either of the authority components are set. The conditions associated with setting any authority component with the two input arguments are tested, and if they do not hold, then an error is raised with the input arguments as irritants.

In update-iri-user and update-iri-host, if obj is not #f and is a string, then it will be parsed for the respective component. Any character which cannot be contained within the component will be encoded as a sequence of percent-encoded bytes. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised. In update-iri-port, if obj is not #f and is a non-negative integer, then it will be set-as is, otherwise #f. An assertion violation is raised with obj as irritant if it is not of the expected type just discussed or #f.

(srfi :275 iri)
procedure
(update-iri-authority iri #f)
(update-iri-authority iri user host port)

Combination procedure which subsumes the above procedures. If a single argument #f is given, then all three components will be set to #f providing this is not in conflict with path. If three arguments are given, then error behaviour is identical to the component-specific setters except that the check for conflict with the existing path is that setting all three to would not be in conflict with the path shape.

(srfi :275 iri)
procedure
(update-iri-query iri obj)
(update-iri-fragment iri obj)

Purely-functional setters for the query and fragment components of an IRI. Error behaviour is identical to update-uri-user, except that there is no check that these are in conflict with the path shape.

(srfi :275 iri)
procedure
(update-iri-path iri obj)

Update the path component of IRI iri with obj. If obj is #f, then the path to be set is the empty path. If obj is a path object, then any character not permissible within a valid IRI path is escaped as percent-encoded bytes corresponding to the charcter's UTF-8 octets. If obj is a string, then it is converted to a path object using a procedure equivalent to string->path, with any illegal characters again escaped as percent-encoded bytes. Next, the conditions associated with setting path are tested for iri and the path object to be set, and if they do not hold, then an error is raised with both input arguments as irritants. An assertion violation is raised with obj as irritant if it is not a path object, a string, or #f.

(srfi :275 iri)
procedure
(iri-equal? iri1 iri2)

Returns #t if the iri1 and iri2 have fields which are equal. A non-relative IRI is not equal to a relative reference and vice versa as relative references have no scheme component. For components other than path and port, two components are only equal if they encode the same sequence of characters and percent-encoded bytes exactly. Two port components are equal if eq? holds between them, and two path componets are equal if path-equal? holds. Returns #f otherwise.

URI programming interface

(srfi :275 uri)
procedure
(uri? uri)
(non-relative-uri? uri)
(relative-uri? uri)
(rfc-uri-reference? uri)
(rfc-absolute-uri? uri)
(empty-uri)
(uri-scheme uri)
(uri-user uri)
(uri-host uri)
(uri-query uri)
(uri-fragment uri)
(uri-port uri)
(uri-path uri)
(uri-authority uri)
(uri-username+password uri)
(uri-path-string uri)
(update-uri-scheme uri obj)
(update-uri-user uri obj)
(update-uri-host uri obj)
(update-uri-port uri obj)
(update-uri-path uri obj)
(update-uri-query uri obj)
(update-uri-fragment uri obj)
(update-uri-authority uri #f)
(update-uri-authority uri user host port)
(uri-equal? uri1 uri2)

Identical programming interfaces to that of IRIs are given for URIs, with argument restriction to the respective URI types, and with the repertoire of characters escaped by setters extended to characters permissible in a given IRI component, but not the equivalent URI component.

Relative reference resolution

(srfi :275 normalize)
procedure
(resolve-iri base iri)

Relative reference resolution of iri against base IRI base. This procedure is purely-functional as it (usually) involves transforming a relative reference into an IRI. An assertion violation is raised with irritant base if it is not a non-relative IRI. Simple examples derived from the RDF Turtle test cases (see test suite section):

(define string-cases (list "g:h" "g" "./g" "g/" "/g" "//g"))

(define base01 (string->iri "http://a/bb/ccc/d;p?q"))
(define base02 (string->iri "http://a/bb/ccc/d/"))
(define base07 (string->iri "file:///a/bb/ccc/d;p?q"))

(define (resolve-with base-iri) ;; Higher-order.  Returns a function.
  (lambda (ref)
    (resolve-iri-reference base-iri (string->iri ref))))

(map (resolve-with base01) string-cases)
→ (list <g:h>
        <http://a/bb/ccc/g> <http://a/bb/ccc/g> <http://a/bb/ccc/g/>
        <http://a/g> <http://g>)

(map (resolve-with base02) string-cases)
→ (list <g:h>
        <http://a/bb/ccc/d/g> <http://a/bb/ccc/d/g> <http://a/bb/ccc/d/g/>
        <http://a/g> <http://g>)

(map (resolve-with base07) string-cases)
→ (list <g:h>
        <file:///a/bb/ccc/g> <file:///a/bb/ccc/g> <file:///a/bb/ccc/g/>
        <file:///g> <file://g>)
  
(srfi :275 normalize)
procedure
(relativize-iri base non-relative-iri)

The inverse of resolve-iri, which elicits a relative reference which is resolveable against base in order to produce non-relative-iri. It is not possible to elicit a path for which rfc-path-abempty? does not hold after calling remove-dot-segments for either argument, and an error is raised with the normalized path and respective argument as irritants. An assertion violation is raised if base is not a non-relative IRI. Note that the argument order is flipped compared to Chicken uri-generic's [6] relative-from procedure.

Examples:

(define base-iri   (string->iri "http://a/bb/ccc/d;p?q"))
(define reference0 (string->iri "g"))
(define resolved0  (string->iri "http://a/bb/ccc/g"))

;; (iri-equal? (relativize base (resolve base v)) v)
(iri-equal? (relativize-iri base-iri (resolve-iri base-iri reference0))
            reference0) → #t

;; (iri-equal? (resolve base (relativize base v)) v)
(iri-equal? (resolve-iri base-iri (relativize-iri base-iri resolved0))
            resolved0)  → #t
(srfi :275 normalize)
procedure
(resolve-uri base uri)
(relativize-uri base non-relative-uri)

Relative reference resolution of an IRI against a base URI, and the inverse, which produces a relative reference given a base IRI and a resolved IRI. Behaviour is structurally identical to that of relative reference resolution for IRIs, and the arguments are instead restricted to non-relative URIs and relative references respectively.

Normalization

The following procedures are purely-functional and structure preserving on the level of components, to avoid some of the problems with string representations being parsed differently post-normalization.

(srfi :275 normalize)
procedure
(normalize-iri-case iri)
(normalize-uri-case uri)

Normalize case-insensitive components (scheme and host), and convert percent-encoded bytes to canonical upper-case.

(srfi :275 normalize)
procedure
(normalize-iri-escape iri)

Normalize an IRI’s escapes, potentially interpreting percent-encoded bytes if part of valid UTF-8 octet sequences as characters. Conversely, characters which are not permissible within an URI component will appear as series of percent-encoded bytes corresponding to UTF-8 octets. This procedure is idempotent: a fully normalized IRI will be normalized to itself, and iri-equal? will hold true between the two.

(srfi :275 normalize)
procedure
(normalize-uri-escape uri)

Normalize an URI’s escapes, potentially interpreting percent-encoded bytes as octets in U.S. ASCII. Conversely, characters which are not permissible within an URI component will appear as series of percent-encoded bytes corresponding to UTF-8 octets.

(srfi :275 normalize)
procedure
(normalize-iri-path iri)
(normalize-uri-path uri)

These procedures normalize path segments given the control segments ‘.’ and ‘..’. While path segments of IRI or URI-relative references are not normalized, these procedures simply have no effect and no error is signalled, consistent with the other normalization procedures. Although it is relatively uncommon for relative paths to appear in non-relative IRIs or URIs, they are permissible e.g. within URN components.

(srfi :275 normalize)
procedure
(normalize-iri iri)
(normalize-uri uri)

These procedures are essentially a sequence of the three normalization procedures described previously.

Equivalence and conversion

(srfi :275 utils)
procedure
(encode-string string)
(encode-string string char-set)

Process string, escaping any characters in the optional character set char-set, which defaults to characters outside of the IRI range (complement of char-set:iri-unreserved). Notably, percent-encoded bytes must not be interpreted as escapes, with the percent-sign treated as a character, and always percent-encoded in turn (as %25).

(srfi :275 utils)
procedure
(decode-string string)
(decode-string string char-set)

Process string, decoding any sequences of percent-encoded bytes as the corresponding UTF-8 octets for characters in the optional character set char-set, which defaults to every character (SRFI 14 char-set:full). When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error is raised with the #\% character and the start of the sequence in which the invalid sub-sequence occured as irritants.

(srfi :275 iri)
procedure
(string->iri string)
(string->non-relative-iri string)
(string->rfc-iri-reference string)
(string->rfc-absolute-iri string)

string->iri parses string into an IRI. An error is raised with the failed character and its position in the string, in the following cases. First, this may occur when a character is not permitted within a given IRI component (unlike the string setters, there is no automatic escaping). Second, when a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised. Finally, if the end of the string was prematurely encountered, then the irritants are the EOF object and the string length.

string->non-relative-iri is a restriction on string->iri in that the scheme component is required. An error is raised with the EOF object and the string length as irritants if the scheme does not end with a colon, and with an illegal character and its position as irritants if this is contained within the scheme. Once a valid scheme has been parsed, the parse continues with the same behaviour as string->iri. This definition of a non-relative IRI (requirement of a scheme component) differs from the RFC 3987 absolute-IRI definition in that a fragment may appear, but we offer this procedure for convenience as many applications (like the RDF N-Triples concrete syntax) forbid relative references but permit non-relative IRIs to have fragment components.

Definitions derived from RFC 3987 verbatim are string->rfc-iri-reference and string->rfc-absolute-iri. The former procedure behaves exactly like string->iri but the naming reflects RFC 3987's IRI-reference terminology. The latter procedure has the same error behaviour as string->non-relative-iri except that the fragment component is strictly forbidden. That is, if the hash character is parsed, then an error is raised with the hash character and its position in the string as irritants.

(srfi :275 iri)
procedure
(iri->string iri)

Serialisation of an IRI as a string. Examples:

(define example-A (string->iri "http://example.org/some/where/place"))
(define example-B (string->iri "urn:/some/where/place"))

(iri->string example-A) → "http://example.org/some/where/place"
(iri->string example-B) → "urn:/some/where/place"
  
(srfi :275 iri)
procedure
(get-iri port)
(get-iri port obj)
(get-non-relative-iri port)
(get-non-relative-iri port obj)
(get-rfc-iri-reference port)
(get-rfc-iri-reference port obj)
(get-rfc-absolute-iri port)
(get-rfc-absolute-iri port obj)

These procedures are the equivalent of the string parsing procedures immediately above, but which parse a stream of text from a port as an IRI. These procedures stop if the port produces a character or the EOF object which is eq? with optional datum obj, which defaults to the EOF object. These low-level procedures are offered because they can be used directly in streaming parsers for data which contains IRIs, such as JSON-LD. All of these procedures raise an assertion violation if port is not a textual input port, and the parse behaviour is identical to the equivalent string procedures except that the position irritant is the equvialent port position.

(srfi :275 uri)
procedure
(string->uri string)
(string->non-relative-uri string)
(string->rfc-uri-reference string)
(string->rfc-absolute-uri string)
(uri->string uri)
(get-uri port)
(get-uri port obj)
(get-non-relative-uri port)
(get-non-relative-uri port obj)
(get-rfc-uri-reference port)
(get-rfc-uri-reference port obj)
(get-rfc-absolute-uri port)
(get-rfc-absolute-uri port obj)

Identical programming interface to that of IRIs are given for URIs, with argument restriction to the respective URI types, and with the repertoire of characters permissible unescaped restricted to those of URIs.

(srfi :275 normalize)
procedure
(iri-eqv? iri1 iri2)
(uri-eqv? uri1 uri2)

Two IRIs are equivalent if iri-equal? holds, or that it holds post-normalization with normalize-iri. Similarly, two URIs are equivalent if uri-equal? holds, or that it holds post-normalization with normalize-uri.

(srfi :275 normalize)
procedure
(iri->uri iri)

Convert an IRI to an URI. This proceeds by encoding any character within certain ranges (see RFC 3987ucschar and iprivate) to a series of percent-encoded bytes corresponding to the character's octets in UTF-8. This URI is also a valid IRI (albeit not normalized) as all URIs are valid IRIs. This procedure is structure-preserving: (non-relative) IRIs are never transformed into relative references, or vice-versa.

(srfi :275 normalize)
procedure
(uri->iri uri)

Convert an URI to an IRI. This procedure can be viewed as upgrading the URI structure to that of an IRI, then normalizing it as an IRI.

Basic paths

This section descibes the basic path sub-library, consisting of two path variants with a common internal segment structure, tagged with whether there is a leading slash. In addition to predicates to discinguish these two path types, we describe predicates which exactly follow the RFC 3986 and 3987 definitions, which are prefixed with rfc-. These do not check that the character set of the path is consistent with RFC 3986 or 3987, at most checking whether the first segment (if any) is empty or contains a colon character. Additionally, the rfc- procedures do not raise an error on non-path objects, instead returning #f, consistent with the generic absolute-path? or relative-path? procedures.

(srfi :275 path)
procedure
(relative-path? obj)
(absolute-path? obj)
(path? obj)

Predicates for relative and absolute path objects, which share a common internal segment structure, returning #t for the respective path objects. path? returns #t if its argument is a relative or absolute path. Returns #f otherwise.

(srfi :275 path)
procedure
(empty-path? obj)

Returns #t if a path is both relative and has no segments. An absolute path with no segments is not considered empty, as it starts with a slash and is represented by the string "/". Returns #f otherwise.

(srfi :275 path)
procedure
(rfc-path-rootless? obj)

Returns #t if the input argument is a relative path with at least one, non-empty segment. Returns #f otherwise.

(srfi :275 path)
procedure
(rfc-path-noscheme? obj)

Returns #t if the input argument is a relative path with at least one, non-empty segment, with the additional condition that the initial segment does not contain a plain (unescaped) colon character. Returns #f otherwise.

(srfi :275 path)
procedure
(rfc-path-abempty? obj)

Returns #t providing that either absolute-path? or empty-path?. Returns #f otherwise.

(srfi :275 path)
procedure
(rfc-path-absolute? obj)

Returns #t providing the first segment is non-empty (i.e. the path does not start with //). Unlike rfc-path-abempty?, this procedure does not hold true for empty paths. Returns #f otherwise.

(srfi :275 path)
procedure
(rfc-path-empty? obj)

Wholly identical to empty-path?.

Examples for the above predicates:

(define rfc-abempty-example  (string->path "//some/where/place"))
(define rfc-absolute-example (string->path "/some/where/place"))
(define rfc-rootless-example (string->path "some:thing/where/place"))
(define rfc-noscheme-example (string->path "some/where/place"))
(define relative-empty-example (string->path ""))
(define absolute-empty-example (string->path "/"))

(define all
  (list rfc-abempty-example rfc-absolute-example
        rfc-rootless-example rfc-noscheme-example
        relative-empty-example absolute-empty-example))

(map absolute-path? all)      → (#t #t   #f #f   #f #t)
(map relative-path? all)      → (#f #f   #t #t   #t #f)
(map empty-path? all)         → (#f #f   #f #f   #t #f)

(map rfc-path-abempty? all)   → (#t #t   #f #f   #t #t) ; empty path is also rfc-path-abempty?
(map rfc-path-absolute? all)  → (#f #t   #f #f   #f #f) ; zero segments is not rfc-path-absolute?
(map rfc-path-rootless? all)  → (#f #f   #t #t   #f #f) ; empty path is not rfc-path-rootless?
(map rfc-path-noscheme? all)  → (#f #f   #f #t   #f #f)
(map rfc-path-empty? all)     → (#f #f   #f #f   #t #f)
(srfi :275 path)
procedure
(empty-relative-path)
(empty-absolute-path)

Helper procedures which take no arguments and construct a relative path or an absolute path with no segments, i.e. a relative path like "" or an absolute path like "/"

(srfi :275 path)
procedure
(path-length path)

Retrieve the length of the path segments, a non-negative integer.

(srfi :275 path)
procedure
(path-segments path)

Serialise the internal segment encoding as a vector of (string) segments. The returned vector cannot be guaranteed to correspond to the internal structure, as there may be percent-encoded bytes requiring a non-string encoding.

(srfi :275 path)
procedure
(path-ref path k)

Retrieve segment of path path at index k as a string.

(srfi :275 path)
procedure
(path-update path k string)

Update segment of path path at index k with string. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised.

(srfi :275 path)
procedure
(path-equal? path1 path2)

Returns #t for two paths path1 and path2, providing that they are both absolute, or are both relative, and that compared pairwise, segments encode the exact same sequence of characters and percent-encoded bytes. Returns #f otherwise.

(srfi :275 path)
procedure
(copy-path path)

Copy the segment structure of a path object and tag it with whether path was relative or absolute.

(srfi :275 path)
procedure
(path->string path)

Serialise path as a string, with no decoding of percent-encoded bytes.

(srfi :275 path)
procedure
(path-string? obj)

Returns #t if x is either a path or string, and if a string, that it could represent a path without additional escaping. The specific check is that a character is permissible within an URI or is within the universal character set ranges permissible within an IRI (see char-set:path-allowed. No error is raised if x is not a string, instead #f is returned. This procedure corresponds to the similarly-named Racket procedure.

(srfi :275 path)
procedure
(string->path string)

Parse string as a path. Both relative paths with a colon in the initial segment, and absolute paths with an empty initial segment, are permitted. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised.

(srfi :275 path)
procedure
(string->relative-path string)
(string->relative-path string noscheme?)

Convert string to a relative path. If the optional noscheme? argument is not #f, and the first segment contains a colon character, then the path output will be prefixed with ./. An error is raised with the string as irritant if it has a leading slash. Error behaviour is otherwise identical to string->path.

(srfi :275 path)
procedure
(string->absolute-path string)
(string->absolute-path string nonempty-ini?)

Convert a string to an absolute path. If the optional nonempt-ini? argument is not #f, and the initial segment is empty (leading double slash), then an error is raised with irritant string. Error behaviour is otherwise identical to string->path.

(srfi :275 path)
procedure
(string->rfc-path-rootless string)
(string->rfc-path-noscheme string)

Helper procedures in which the final optional argument to string->relative-path is set to #f and #t respectively.

(srfi :275 path)
procedure
(string->rfc-path-abempty string)
(string->rfc-path-absolute string)

Helper procedures in which the final optional argument to string->absolute-path is set to #f and #t respectively.

(srfi :275 path)
procedure
(vector->relative-path vector)
(vector->relative-path vector noscheme?)

Convert a vector of string segments to a relative path. This procedure behaves similarly to string->relative-path with the exception that that slashes are always escaped, so a slash at the start of the first segment does not raise an error. An error is raised if a segment is not a string, with the segment, its position, and vector as irritants, and if the initial segment is empty (serialised as leading slash), then an error is raised with the (string) segment, its position, and vector as irritants.

(srfi :275 path)
procedure
(vector->absolute-path vector)
(vector->absolute-path vector nonempty-ini?)

Convert a vector of string segments to an absolute path. This procedure behaves similarly to string->absolute-path with the exception that that slashes are always escaped, so a slash at the start of the first segment does not raise an error. An error is raised if a segment is not a string, with the segment, its position, and vector as irritants.

(srfi :275 path)
procedure
(vector->rfc-path-rootless vector)
(vector->rfc-path-noscheme vector)

Helper procedures in which the final optional argument to vector->relative-path is set to #f and #t respectively.

(srfi :275 path)
procedure
(vector->rfc-path-abempty vec)
(vector->rfc-path-absolute vec)

Helper procedures in which the final optional argument to vector->absolute-path is set to #f and #t respectively.

(srfi :275 path)
procedure
(build-path)
(build-path base string1 ...)

Variadic procedure to build a path from a base path and any number of additional string segments. If no arguments are given, then the empty path is returned. This procedure corresponds to the similarly-named Racket procedure.

If base is a string, then the procedure behaves like vector->relative-path. If base is a path, then the procedure concatenates the base path with a relative path corresponding to the remaining segments. This is accomplished using the merge-paths procedure described in RFC 3986 section 5.2.3. The slash character is always escaped within segments.

By admitting the base as potentially another path, this allows one to build an absolute path from a list of segments, by choosing (empty-absolute-path) as base.

An error is raised if a segment is not a string, with the segment, and its position, as irritants. Additionally, if the initial segment (after expanding the base path) is empty, then an error is raised with the (string) segment and its position as irritants.

(srfi :275 path)
procedure
(remove-dot-segments path)

Remove dotted segments (‘.’ and ‘..’) in path by interpreting them alongside the path segment structure. This procedure corresponds to RFC 3986 section 5.2.4.

(srfi :275 path)
procedure
(merge-paths base path)
(merge-paths base path base-has-authority?)

Merge base path base with path. As per RFC 3986 section 5.2.3, if base is empty and it has an authority component, then path is concatenated with the reference path. This is controlled by the whether the optional base-has-authority? argument is non-#f, and is mostly important during relative reference resolution. This defaults to #f as this procedure in isolation assumes no authority component.

(srfi :275 iri)
procedure
(iri-path-absolute? iri)
(iri-path-relative? iri)
(iri-path-empty? iri)
(iri-path-rfc-abempty? iri)
(iri-path-rfc-absolute? iri)
(iri-path-rfc-rootless? iri)
(iri-path-rfc-noscheme? iri)
(iri-path-rfc-empty? iri)

Helper procedures which call the path library’s corresponding procedures for path shape of an IRI. Unlike the equvialent path procedures, these raise an assertion violation with iri as irritant if it is not an IRI. Examples:

(define rfc-abempty-example  (string->iri "urn://myhost//some/where/place"))
(define rfc-absolute-example (string->iri "urn://myhost/some/where/place"))
(define rfc-rootless-example (string->iri "urn:some:thing/where/place"))
(define rfc-noscheme-example (string->iri "urn:some/where/place"))
(define relative-empty-example (string->iri "urn:"))
(define absolute-empty-example (string->iri "urn:/"))

(define all
  (list rfc-abempty-example rfc-absolute-example
        rfc-rootless-example rfc-noscheme-example
        relative-empty-example absolute-empty-example))

(map iri-path-string all)
→ ("//some/where/place"     "/some/where/place"
    "some:thing/where/place" "some/where/place"
    ""                       "/")

(map iri-path-absolute? all)      → (#t #t   #f #f   #f #t)
(map iri-path-relative? all)      → (#f #f   #t #t   #t #f)
(map iri-path-empty? all)         → (#f #f   #f #f   #t #f)

(map iri-path-rfc-abempty? all)   → (#t #t   #f #f   #t #t) ; empty path is also rfc-path-abempty?
(map iri-path-rfc-absolute? all)  → (#f #t   #f #f   #f #f) ; zero segments is not rfc-path-absolute?
(map iri-path-rfc-rootless? all)  → (#f #f   #t #t   #f #f) ; empty path is not rfc-path-rootless?
(map iri-path-rfc-noscheme? all)  → (#f #f   #f #t   #f #f)
(map iri-path-rfc-empty? all)     → (#f #f   #f #f   #t #f)
  
(srfi :275 iri)
procedure
(iri-path-segments ident)
(iri-path-segment-ref ident k)
(iri-path-segment-update ident k str)

Procedures which effectively wrap path-segments, path-ref and path-update respectively. These are offered to expose path internals without requiring to import the path sub-library. iri-path-segment-update has additional error behaviour beyond path-update: characters not permissible within an IRI path segment will be automatically escaped as sequences of percent-encoded bytes, comparable to the behaviour of update-iri-path on segments.

(srfi :275 uri)
procedure
(uri-path-absolute? uri)
(uri-path-relative? uri)
(uri-path-empty? uri)
(uri-path-rfc-abempty? uri)
(uri-path-rfc-absolute? uri)
(uri-path-rfc-rootless? uri)
(uri-path-rfc-noscheme? uri)
(uri-path-rfc-empty? uri)

Helper procedures which call the path library’s corresponding procedures for path shape of an URI. Unlike the equvialent path procedures, these raise an assertion violation with ident as irritant if it is not an URI.

(srfi :275 uri)
procedure
(uri-path-segments uri)
(uri-path-segment-ref uri k)
(uri-path-segment-update uri k str)

Procedures which effectively wrap path-segments, path-ref and path-update respectively. In iri-path-segment-update, characters not permissible within an URI path segment will be automatically escaped.

Miscellaneous utilities

(srfi :275 utils)
procedure
(utf8->string/raise bytevector)

R6RS’s utf8->string procedure silently substitutes the unicode replacement character (U+FFFD) given an invalid UTF-8 sequence. If the information is available, then this procedure raises an assertion violation with the offending octet and its position within the bytevector as irritants. Otherwise, the irritants are the offending octet alone.

(srfi :275 utils)
procedure
(percent-encoding->string k)

Represent octet k (representing a percent-encoded byte) as a string in its canonical, upper-case form % 0-F 0-F. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256. Examples:

(percent-encoding->string #x40) → "%40"
(percent-encoding->string #xFA) → "%FA"
(percent-encoding->string #x00) → "%00"

(percent-encoding->string -12.0  → ERROR)
(percent-encoding->string #x100) → ERROR)
  
(srfi :275 utils)
procedure
(percent-encoding->u8-list pct)

Represent octet k (reperesnting a percent-encoded byte) as a list of UTF-8 octets corresponding to the canonical string representation of the percent-encoded byte. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256. Examples:

(percent-encoding->u8-list #x40) → (#x25 #x34 #x30)
(percent-encoding->u8-list #xFA) → (#x25 #x46 #x41)
(percent-encoding->u8-list #x00) → (#x25 #x30 #x30)

(map integer->char (percent-encoding->u8-list #xFA)) → (#\% #\F #\A)
  
(srfi :275 utils)
procedure
(percent-encoding->reverse-u8-list pct)

Represent octet k (reperesnting a percent-encoded byte) as a list of UTF-8 octets corresponding to the canonical string representation of the percent-encoded byte, this time in reverse order. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256.

(srfi :275 utils)
procedure
(username+password str)

Split strings retrieved from an IRI or an URI’s user component into username and password components. Again, a colon which is escaped is not considered a delimiter between username and password. Examples:

(username+password (iri-user (string->iri"//foo:bar:qux@host")))   → "foo"        "bar:qux"
(username+password (iri-user (string->iri"//foo%3Abar:qux@host"))) → "foo%3Abar"  "qux"
(username+password (iri-user (string->iri"//@host")))     → ""    #f
(username+password (iri-user (string->iri "//")))         → #f    #f
(username+password (iri-user (string->iri "//foo@host"))) → "foo" #f
  
(srfi :275 utils)
procedure
(user-settable? ident obj)
(host-settable? ident obj)
(port-settable? ident obj)
(path-settable? ident obj)

Helper procedures which hold providing that updating the respective component of IRI or URI ident to obj is legal. This directly reflects the RFC 3986 and 3987 ABNF, but it’s convenient to provide a single predicate as the conditions to check are somewhat elaborate with respect to path shape. These conditions are described in the error behaviour of setters section. For efficiency, obj is not parsed and the only check is that it is not #f. This means that subsequent procedures called on these arguments may error. An assertion violation is raised if the first argument is not an IRI or URI, with that argument as irritant.

Character sets

Note that (srfi :275 iri) and (srfi :275 uri) export some of the same identifiers where the set of characters is the same in RFC 3986 and 3987: e.g. char-set:reserved. Character sets are named for the RFC 3986 and 3987 ABNF productions, e.g. char-set:uri-userinfo.

(srfi :275 uri)
datum
char-set:uri-unreserved
char-set:uri-allowed

RFC 3986 unreserved character set, and the repertoire of all unescaped characters permissible in an URI.

(srfi :275 uri)
datum
char-set:gen-delims
char-set:sub-delims
char-set:reserved

RFC 3986 gen-delims and sub-delims character sets, and their union, the reserved range. These character sets are also exported by the (srfi :275 iri) library for convenience.

(srfi :275 uri)
datum
char-set:scheme
char-set:uri-userinfo
char-set:uri-reg-name
char-set:uri-segment
char-set:uri-query
char-set:uri-fragment

RFC 3986 component-specific character sets. In RFC 3986, but not RFC 3987, query and fragment share the same repertoire.

(srfi :275 iri)
datum
char-set:iri-private
char-set:ucschar
char-set:iri-unreserved
char-set:iri-allowed

RFC 3987 iunreserved character set, which is the union of char-set:uri-unreserved and most of the universal character set. The permissible ranges in the UCS are exported as char-set:ucschar, and the additional char-set:iri-private range exports the additional characters permissible in RFC 3987 query components. Finally, char-set:iri-allowed is the repertoire of all characters permissible in an IRI.

(srfi :275 iri)
datum
char-set:scheme
char-set:iri-userinfo
char-set:iri-reg-name
char-set:iri-segment
char-set:iri-query
char-set:iri-fragment

RFC 3987 component-specific character sets. The scheme component is the same as in RFC 3986. Query is more expansive in RFC 3987 and includes the char-set:iri-private range, unlike fragment.

(srfi :275 path)
datum
char-set:path-allowed

Union of char-set:uri-allowed and char-set:ucschar described above. While this is minimally restritive, it does contain (a few) characters disallowed within IRI segments. However, any illegal characters will be escaped in the path setters for URIs and IRIs.

Test suite

In this section, we describe various test cases and the specific behaviour they evaluate. Because RFC 3986 and 3987 only provide examples for relative reference resolution, it is important to specify the exact behaviour of a correct implementation, especially with respect to normalization. It is expected that these test cases could be the basis of a more comprehensive property-based test suite, e.g. test the entire range of reserved and unreserved characters in a particular URI component.

URI case normalization test cases (normalize-uri-case)

All lower-case preserved
<http://example.org/ex#test><http://example.org/ex#test>
Mixed-case scheme to lower-case
<HttP://example.org/ex#test><http://example.org/ex#test>
Mixed-case user preserved
<http://MySelf@example.org/Examp#test><http://MySelf@example.org/Examp#test>
Mixed-case host to lower-case
<http://Example.ORG/ex#test><http://example.org/ex#test>
Mixed case path preserved
<http://example.org/Examp#test><http://example.org/Examp#test>
Mixed case query preserved
<http://example.org/examp?Qua#test><http://example.org/examp?Qua#test>
Mixed case fragment preserved
<http://example.org/examp#TeSt><http://example.org/examp#TeSt>
User percent-encodings to upper-case
<http://%aA@%AA%AB%AC%AD%AE/some/where/place><http://%AA@%AA%AB%AC%AD%AE/some/where/place>
Host percent-encodings to upper-case
<http://%aa%Ab%AC%aD%AE/some/where/place><http://%AA%AB%AC%AD%AE/some/where/place>
Path percent-encodings to upper-case
<http://myname@example.org/%Fa/%FB/%fC><http://myname@example.org/%FA/%FB/%FC>
Query percent-encodings to upper-case
<http://myname@example.org/%FA/%FB/%FC?%ff><http://myname@example.org/%FA/%FB/%FC?%FF>
Fragment percent-encodings to upper-case
<http://myname@example.org/%FA/%FB/%FC#%ff><http://myname@example.org/%FA/%FB/%FC#%FF>

IRI case normalization test cases (normalize-iri-case)

Host non-U.S. ASCII is case-sensitive
<http://CRÊPES.example.org><http://crÊpes.example.org> No equivalent for scheme as that only contains ASCII even in IRIs

URI escape (percent-encoded byte) normalization test cases (normalize-uri-escape)

User reserved character not escaped
<http://my!name@example.org/ex#test><http://my!name@example.org/ex#test>
Host reserved character not escaped
<http://myname@!example.org/ex#test><http://myname@!example.org/ex#test>
Path reserved character not escaped
<http://myname@example.org/ex!#test><http://myname@example.org/ex!#test>
Query reserved character not escaped
<http://myname@example.org/ex?!a#test><http://myname@example.org/ex?!a#test>
Fragment reserved character not escaped
<http://myname@example.org/ex?a#!test><http://myname@example.org/ex?a#!test>
User reserved escape not decoded
<http://my%40name@example.org/ex#test><http://my%40name@example.org/ex#test>
Host reserved escape not decoded
<http://myname@ex%40ample.org/ex#test><http://myname@ex%40ample.org/ex#test>
Path reserved escape not decoded
<http://myname@example.org/e%40x#test><http://myname@example.org/e%40x#test>
Query reserved escape not decoded
<http://myname@example.org/ex?a%40#test><http://myname@example.org/ex?a%40#test>
Fragment reserved escape not decoded
<http://myname@example.org/ex?a#t%40est><http://myname@example.org/ex?a#t%40est>
User permissible escape decoded
<http://my%2Ename@example.org/ex?a#test><http://my.name@example.org/ex?a#test>
Host permissible escape decoded
<http://myname@example%2Eorg/ex?a#test><http://myname@example.org/ex?a#test>
Path permissible escape decoded
<http://myname@example.org/misc%2Etxt#test><http://myname@example.org/misc.txt#test>
Query permissible escape decoded
<http://myname@example.org/misc.txt?%2E%2E%2E><http://myname@example.org/misc.txt?...>
Fragment permissible escape decoded
<http://myname@example.org/misc.txt#line%31%30><http://myname@example.org/misc.txt#line10>
User illegal characters are escaped
<http://dosh£@crepes.example.org><http://dosh%C2%A3@crepes.example.org>
Host illegal characters are escaped
<http://crêpes.example.org><http://cr%C3%AApes.example.org>
Path illegal characters are escaped
<http://crepes.example.org/in/Rhône><http://crepes.example.org/in/Rh%C3%B4ne>
Query illegal characters are escaped
<http://crepes.example.org/in/Rennes?Dim.‥Sam.><http://crepes.example.org/in/Rennes?Dim.%E2%80%A5Sam.>
Fragment illegal characters are escaped
<http://crepes.example.org/in/Rennes#L'Étage><http://crepes.example.org/in/Rennes#L'%C3%89tage>

IRI escape normalization test cases (normalize-iri-escape)

User permissible escape decoded
<http://dosh%C2%A3@crepes.example.org><http://dosh£@crepes.example.org>
Host permissible escape decoded
<http://cr%C3%AApes.example.org><http://crêpes.example.org>
Path illegal characters are escaped
<http://crepes.example.org/in/Rh%C3%B4ne><http://crepes.example.org/in/Rhône>
Query illegal characters are escaped
<http://crepes.example.org/in/Rennes?Dim.%E2%80%A5Sam.><http://crepes.example.org/in/Rennes?Dim.‥Sam.>
Fragment illegal characters are escaped
<http://crepes.example.org/in/Rennes#L'%C3%89tage><http://crepes.example.org/in/Rennes#L'Étage>
Wholly escaped wholly decoded
<https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82><https://en.wiktionary.org/wiki/Ῥόδος>
Partially normalized wholly decoded
<https://example.org/music/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh><https://example.org/music/Éirigh'sCuirOrtDoChuidÉadaigh>
Wholly normalized not reencoded
<https://en.wiktionary.org/wiki/Ῥόδος><https://en.wiktionary.org/wiki/Ῥόδος>
Repeated normalization (idempotence)
<https://en.wiktionary.org/wiki/Ῥόδος><https://en.wiktionary.org/wiki/Ῥόδος> In this test case, normalize-iri-escape should be called multiple times, i.e. (compose normalize-iri-escape normalize-iri-escape)

Path normalization test cases (normalize-uri-path and normalize-iri-path)

Absolute path without dotted segments unchanged
<http://example.org/some/where/place><http://example.org/some/where/place>
No authority absolute path without dotted segments unchanged
<urn:/some/where/place><urn:/some/where/place>
Relative path without dotted segments unchanged
<urn:some/where/place><urn:some/where/place>
Absolute path eliminates single dotted segments
<urn:/some/./where/././place/./><urn:/some/where/place/>
Relative path eliminates single dotted segments
<urn:some/./where/././place/./><urn:some/where/place/>
Absolute path empty segments treated like non-empty
<urn:/some//where//place//><urn:/some//where//place//>
Relative path empty segments treated like non-empty
<urn:some//where//place//><urn:some//where//place//>
Single leading slash not normalized
</></>
Multiple leading slashes not normalized
<//><//>
Absolute path reference not normalized (double-dot)
</a/b/../../c></a/b/../../c>
Absolute path reference not normalized (single-dot)
</a/b/././c></a/b/././c>
Absolute path reference not normalized (mixed dotted)
</a/b/../c/././d></a/b/../c/././d>
Relative path reference not normalized (double-dot)
<a/b/../../c><a/b/../../c>
Relative path reference not normalized (single-dot)
<a/b/././c><a/b/././c>
Relative path reference not normalized (mixed dotted)
<a/b/../c/././d><a/b/../c/././d>
1-segment path reference with leading dot not normalized
<./def><./def>
1-segment path reference with leading dot not normalized (colon)
<./abc:def><./abc:def>
Additional relative path case not normalized
<../../abc/./def><../../abc/./def>
Relative path must not be normalized to an absolute path
<foo:a/b/../.././../../e><foo:e> From Haskell network-uri [5]
Empty segments eliminated like non-empty (1)
<http://example.com////../..><http://example.com//> From Webkit [13]
Empty segments eliminated like non-empty (2)
<http://example.com/foo/bar//../..><http://example.com/foo/> From Webkit [13]
Empty segments eliminated like non-empty (3)
<http://example.com/foo/bar//..><http://example.com/foo/bar/> From Webkit [13]
General test case 1
<http://example/a/b/../../c><http://example/c> From Haskell network-uri [5]
General test case 2
<http://example/a/b/c/../../><http://example/a/> From Haskell network-uri [5]
General test case 3
<http://example/a/b/c/./><http://example/a/b/c/> From Haskell network-uri [5]
General test case 4
<http://example/a/b/c/.././><http://example/a/b/> From Haskell network-uri [5]
General test case 5
<http://example/a/b/c/d/../../../../e><http://example/e> From Haskell network-uri [5]
General test case 6
<http://example/a/b/c/d/../.././../../e><http://example/e> From Haskell network-uri [5]
General test case 7
<http://example/a/b/../.././../../e><http://example/e> From Haskell network-uri [5]

URI to IRI conversion (uri->iri)

Wholly escaped path normalized
<https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82><https://en.wiktionary.org/wiki/Ῥόδος>
Partially escaped path normalized
<https://example.org/ceol/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh><https://example.org/ceol/Éirigh'sCuirOrtDoChuidÉadaigh>

IRI to URI conversion (iri->uri)

Wholly escaped path normalized
<https://en.wiktionary.org/wiki/Ῥόδος><https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82>
Partially escapable path escaped
<https://example.org/ceol/Éirigh'sCuirOrtDoChuidÉadaigh><https://example.org/ceol/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh>

Relative reference resolution (resolve-iri and resolve-uri)

The following RDF 1.1 Turtle test cases for IRI resolution must succeed. These tests are based on those of RFC 3986, and are structured as triples in which the subject and predicate components are absolute IRIs (URNs), but where the object component is a relative IRI. A base IRI is declared at the top of the Turtle (.ttl) file, and the corresponding N-Triples (.nt) file lists the resolved triples. In the SRFI 275 sample implementation, we inline these tests, with the objects extracted and resolved against the base IRI explicitly.

  1. IRI-resolution-01.ttlIRI-resolution-01.nt    (with base IRI <http://a/bb/ccc/d;p?q>)
  2. IRI-resolution-02.ttlIRI-resolution-02.nt    (with base IRI <http://a/bb/ccc/d/>)
  3. IRI-resolution-07.ttlIRI-resolution-07.nt    (with base IRI <file:///a/bb/ccc/d;p?q>)
  4. IRI-resolution-08.ttlIRI-resolution-08.nt

We require the following abnormal test cases from Chicken uri-generic [6], when resolved against base IRI <http://a/b/c/d;p?q>, to yield the following results:

Over-deep ‘..’ traversal is clamped at root, not an error
<../../../g>  →  <http://a/g>
<../../../../g>  →  <http://a/g>
<../../../..>  →  <http://a/>
<../../../../>  →  <http://a/>
Dotted segments in absolute paths are removed
</./g>  →  <http://a/g>
</../g>  →  <http://a/g>
These are not dot-segments; literal names, not traversal
<g..>  →  <http://a/b/c/g..>
<..g>  →  <http://a/b/c/..g>
Dots do not affect query or fragment
<g?y/./x>  →  <http://a/b/c/g?y/./x>
<g?y/../x>  →  <http://a/b/c/g?y/../x>
<g#s/./x>  →  <http://a/b/c/g#s/./x>
<g#s/../x>  →  <http://a/b/c/g#s/../x>

Reference resolution must ensure that a path is never elicited which would be mistaken for the double slash after a scheme’s colon. (The selected strategy is to prefix a ./ to the path, making it relative, but in principle prepending a slash to the segments i.e. after the initial slash, /.///, is equivalent.)

Base IRI <f:/a>
<.//g>  →  <f:.///g>
Base IRI <f:/a/>
<..//g>  →  <f:.///g>

The following test cases are derived from the SWAP project's [8] uripath.py These tests are formatted here as base URI or IRI, relative reference, and expected result.

Base IRI <foo:xyz>
<bar:abc>  →  <bar:abc>
Base IRI <http://example/x/y/z>
<../abc>  →  <http://example/x/abc>
Base IRI <http://example2/x/y/z>
<//example/x/abc>  →  <http://example/x/abc>
Base IRI <http://ex/x/y/z>
<../r>  →  <http://ex/x/r>
Base IRI <http://ex/x/y>
<q/r>  →  <http://ex/x/q/r>
Base IRI <http://ex/x/y>
<q/r#s>  →  <http://ex/x/q/r#s>
Base IRI <http://ex/x/y>
<q/r#s/t>  →  <http://ex/x/q/r#s/t>
Base IRI <http://ex/x/y>
<ftp://ex/x/q/r>  →  <ftp://ex/x/q/r>
Base IRI <http://ex/x/y>
<y>  →  <http://ex/x/y>
Base IRI <http://ex/x/y/>
<.>  →  <http://ex/x/y/>
Base IRI <http://ex/x/y/pdq>
<pdq>  →  <http://ex/x/y/pdq>
Base IRI <http://ex/x/y/>
<z/>  →  <http://ex/x/y/z/>
Base IRI <file:/swap/test/animal.rdf>
<animal.rdf#Animal>  →  <file:/swap/test/animal.rdf#Animal>
Base IRI <file:/e/x/y/z>
<../abc>  →  <file:/e/x/abc>
Base IRI <file:/example2/x/y/z>
<../../../example/x/abc>  →  <file:/example/x/abc>
Base IRI <file:/ex/x/y/z>
<../r>  →  <file:/ex/x/r>
Base IRI <file:/ex/x/y/z>
<../../../r>  →  <file:/r>
Base IRI <file:/ex/x/y>
<q/r>  →  <file:/ex/x/q/r>
Base IRI <file:/ex/x/y>
<q/r#s>  →  <file:/ex/x/q/r#s>
Base IRI <file:/ex/x/y>
<q/r#>  →  <file:/ex/x/q/r#>
Base IRI <file:/ex/x/y>
<q/r#s/t>  →  <file:/ex/x/q/r#s/t>
Base IRI <file:/ex/x/y>
<ftp://ex/x/q/r>  →  <ftp://ex/x/q/r>
Base IRI <file:/ex/x/y>
<y>  →  <file:/ex/x/y>
Base IRI <file:/ex/x/y/>
<.>  →  <file:/ex/x/y/>
Base IRI <file:/ex/x/y/pdq>
<pdq>  →  <file:/ex/x/y/pdq>
Base IRI <file:/ex/x/y/>
<z/>  →  <file:/ex/x/y/z/>
Base IRI <file:/devel/WWW/2000/10/swap/test/reluri-1.n3>
<//meetings.example.com/cal#m1>  →  <file://meetings.example.com/cal#m1>
Base IRI <file:/home/connolly/w3ccvs/WWW/2000/10/swap/test/reluri-1.n3>
<//meetings.example.com/cal#m1>  →  <file://meetings.example.com/cal#m1>
Base IRI <file:/some/dir/foo>
<.#blort>  →  <file:/some/dir/#blort>
Base IRI <file:/some/dir/foo>
<.#>  →  <file:/some/dir/#>
Base IRI <http://example/x/y%2Fz> (see here)
<abc>  →  <http://example/x/abc>
Base IRI <http://example/x/y/z> (see here)
<../../x%2Fabc>  →  <http://example/x%2Fabc>
Base IRI <http://example/x/y%2Fz> (see here)
<../x%2Fabc>  →  <http://example/x%2Fabc>
Base IRI <http://example/x%2Fy/z> (see here)
<abc>  →  <http://example/x%2Fy/abc>
Base IRI <http://example/x/abc.efg>
<.>  →  <http://example/x/>

The following test cases are derived from the Haskell network-uri [5] package. These tests are formatted here as base URI or IRI, relative reference, and expected result. The first set of cases are tricky cases of non-relative paths together with query and fragment.

Base IRI <mailto:local1@domain1?query1>
<local2@domain2>  →  <mailto:local2@domain2>
Base IRI <mailto:local1@domain1>
<local2@domain2?query2>  →  <mailto:local2@domain2?query2>
Base IRI <mailto:local1@domain1?query1>
<local2@domain2?query2>  →  <mailto:local2@domain2?query2>
Base IRI <mailto:local@domain?query1>
<?query2>  →  <mailto:local@domain?query2>
Base IRI <mailto:?query1>
<local@domain?query2>  →  <mailto:local@domain?query2>
Base IRI <mailto:local@domain?query1>
<?query2>  →  <mailto:local@domain?query2>
Base IRI <foo:bar>
<http://example/a/b?c/../d>  →  <http://example/a/b?c/../d>
Base IRI <foo:bar>
<http://example/a/b#c/../d>  →  <http://example/a/b#c/../d>

The remaining set of tests from network-uri [5] are with respect to dealing with the final segment when inverting (see next section). See here. These test cases are not exactly invertible, however as they do not all produce the exact same reference due to relative segments.

Base IRI <http://www.example.com/data/limit/..>
<test.xml>  →  <http://www.example.com/data/limit/test.xml>
Base IRI <file:/some/dir/foo>
<./#blort>  →  <file:/some/dir/#blort>
Base IRI <file:/some/dir/foo>
<./#>  →  <file:/some/dir/#>
Base IRI <file:/some/dir/..>
<./#blort>  →  <file:/some/dir/#blort>
Base IRI <http://example.org/base/uri>
<http:this>  →  <http:this>
Base IRI <http:base>
<http:this>  →  <http:this>
Base IRI <f://example.org/base/a>
<b/c//d/e>  →  <f://example.org/base/b/c//d/e>
Base IRI <mid:m@example.ord/c@example.org>
<m2@example.ord/c2@example.org>  →  <mid:m@example.ord/m2@example.ord/c2@example.org>
Base IRI <file:///C:/DEV/Haskell/lib/HXmlToolbox-3.01/examples/>
<mini1.xml>  →  <file:///C:/DEV/Haskell/lib/HXmlToolbox-3.01/examples/mini1.xml>
Base IRI <foo:a/y/z>
<../b/c>  →  <foo:a/b/c>

Relativization with respect to a base IRI (relativize-iri and relativize-uri)

These procedures are effectively the inverse of resolve-iri and resolve-uri. However, not all of the test cases described above succeed verbatim, as e.g. the references may include dotted segments which would be noramlised during the relative reference resolution process. Of the test cases descibed above, we specifically require that the test cases derived from the Python SWAP project as all of these test cases are invertible exactly.

Recall that test data above are formatted as a base URI or IRI, a relative reference, and an expected result of reference resolution. A test of relative reference extraction succeeds if given the base IRI or URI and the expected result of resolution, the relative reference (described in the middle column) is elicited.

The following inversion must hold, but note that we elicit relative reference </g>, which is only the same as previous reference <.//g> after calling remove-dot-segments.

Base IRI <f:/a>
<f:.///g>  →  </g>
Base IRI <f:/a/>
<f:.///g>  →  <..//g>

Finally, the following test cases included with Chicken's uri-generic library [6] must hold for base <http://a/b/c/d;p?q>:

Unusual directory in base, file in target
<http://a/b/c>  →  <../c>
Could in principle elicit </>, but that's not convenient
<http://a/>  →  <../..>
No relative representation possible
<http://a>  →  <//a>
Identical scheme elicits output resolved identifier
<ftp://a/b/c/d;p?q>  →  <ftp://a/b/c/d;p?q>
<ftp://x/y/z;a?b>  →  <ftp://x/y/z;a?b>
General test cases
<http://a/b/c/d;p?q>  →  <d;p>
<http://a/b/c/e>  →  <e>
<http://a/b/c/>  →  <.>
<http://a/b/e>  →  <../e>
<http://a/b/>  →  <..>
<http://b>  →  <//b>
<http://b/>  →  <//b/>
<http://b/c>  →  <//b/c>

UTF-8 interpretation

In this section, we detail negative test cases which must be rejected. These go beyond the RFC 3986 and 3987 specifications which permit percent-encoded bytes which may correspond to invalid UTF-8 octet sequences. No encodings other than UTF-8 are supported. The following example cases are for the URI component. The same sequences must be rejected in every component in which percent-encoded bytes may appear (every component except scheme and port).

Encode all four byte lengths in a single call
(encode-string "a béc♂d😎e" (char-set-complement char-set:uri-allowed))"a%20b%C3%A9c%E2%99%82d%F0%9F%98%8Ee"
Decode everything except ♂ (U+2642, encoded as %E2%99%82)
(decode-string "a%20b%C3%A9c%E2%99%82d%F0%9F%98%8Ee" (char-set-complement char-set:uri-allowed))"a béc%E2%99%82d😎e"
Direct encoding treats percent as a character
(encode-string "foo%20bar" (char-set-difference char-set:full char-set:uri-allowed))"foo%2520bar"
Escaped percent round trip during direct decoding
(decode-string "foo%2520bar" char-set:full)"foo%20bar"

The following strings must be rejected by decode-string (in contrast encode-string does not interpret escapes):

Reject overlong 2-byte encoding of U+0020 (space)
"a%C0%A0b"
Reject overlong 3-byte encoding of U+0020
"a%E0%80%A0b"
Reject overlong 4-byte encoding of U+0020
"a%E0%80%80%A0b"
Reject overlong 3-byte encoding of U+00E9 (é)
"a%E0%80%A9b"
Reject incomplete 2-byte sequence (missing continuation byte)
"a%C0b"
Reject incomplete 3-byte sequence (missing third byte for ♂ U+2642)
"a%E2%99"
Reject incomplete 4-byte sequence (missing fourth byte for 😎 U+1F60E)
"a%F0%9F%98"
Reject non-encoded character where a continuation byte is expected
"a%F0%9F%98x"
Reject UTF-8 continuation byte without a leading byte
"a%A9"

Additionally, we require that these sequences are rejected by the different parsers, as well as setters for IRI and URI components (e.g. update-iri-host). The utf8->string/raise procedure is provided for detecting these sequences in external code.

IP literals

We expect the following edge-cases to have the corresponding hostname verbatim:

All-zeroes / unspecified address
<http://[::]/>"[::]"
Loopback
<http://[::1]/>"[::1]"
Trailing ::
<http://[1::]/">"[1::]"
Trailing ::, prefix
<http://[2001:db8::]/>"[2001:db8::]"
Leading ::
<http://[::2001:db8]/>"[::2001:db8]"
IPv4-mapped
<http://[::ffff:192.0.2.1]/>"[::ffff:192.0.2.1]"
IPv4-translated (RFC 6052)
<http://[64:ff9b::192.0.2.1]/>"[64:ff9b::192.0.2.1]"
Full address with port, path, query and fragment
<http://[2001:db8::1]:8080/path?query#fragment>"[2001:db8::1]"

Negative test cases which must be rejected:

Reject more than one :: compression marker
<http://[2001:db8:::1]/>
Reject nine groups (maximum is eight)
<http://[2001:db8:a:b:c:d:e:f:1]/>
Reject seven groups with no :: (too few)
http://[2001:db8:a:b:c:d:e]/< >
Reject octet group exceeding four hex digits
<http://[2001:db8::fffff]/>
Reject same; leading zeros push group to five digits
<http://[2001:00db8::0001]/>
Reject incomplete IPv4 suffix (three octets, not four)
<http://[2001:db8::192.0.2]/>
Reject two separate :: markers
<http://[2001:db8:a::b::c]/>
Reject only two groups, no ::
<http://[2001:db8]/>
Reject unclosed bracket
<http://[2001:db8>

Parsing expected segments

schemeuserhostportpathqueryfragment
Empty URI: <>
N/A#f#f#f#f#f#f
Empty authority: <//>
N/A#f""#f#f#f#f
Empty user: <//@>
N/A""#f#f#f#f#f
Empty port: <//:>
N/A#f""#f#f#f#f
Empty query: <?>
N/A#f#f#f#f""#f
Empty fragment: <#>
N/A#f#f#f#f#f""
Path which looks like a hostname: <example.org>
N/A#f#f#f"example.org"#f#f
URN-like: <urn:something>
"urn"#f#f#f"something"#f#f
URN-like, path looks like hostname: <urn:example.org>
"urn"#f#f#f"example.org"#f#f
Path which looks like a URN: <./urn:something>
N/A#f#f#f"./urn:something"#f#f
User with colon segment: <http://a:b@c:29>
"http""a:b""c"29#f#f#f
User-like component appears as path: <http::@c:29>
"http"#f#f#f":@c:29"#f#f
Host-like component appears as user: <http://example.org:b@d/>
"http""example.org:b""d"#f"/"#f#f
Padded port as numeric value: <http://example.org:000080>
"http"#f"example.org"80#f#f#f
Query component with question mark: <http://example.org/abcd?efgh?ijkl>
"http"#f"example.org"#f"/abcd""efgh?ijkl"#f
Fragment component with question mark: <http://example.org/abcd#efgh?ijkl>
"http"#f"example.org"#f"/abcd"#f"efgh?ijkl"
Path where first segment looks like host: <http:///some/where/place>
"http"#f""#f"/some/where/place"#f#f
Scheme with nil host: <foo:>
"foo"#f#f#f#f#f#f
Scheme with path, empty host: <foo:////g>
"foo"#f""#f"//g"#f#f
Scheme with path, nil host: <foo:.///g>
"foo"#f#f#f".///g"#f#f
Scheme with non-empty host: <foo://g>
"foo"#f"g"#f#f#f#f
All components filled out: <http://user@example.org:80/some/where/place?qua#ought>
"http""user""example.org"80"/some/where/place""qua""ought"
All components except user filled out: <http://example.org:80/some/where/place?qua#ought>
"http"#f"example.org"80"/some/where/place""qua""ought"
All components except host filled out: <http://user@:80/some/where/place?qua#ought>
"http""user"#f80"/some/where/place""qua""ought"
All components except port filled out: <http://user@example.org/some/where/place?qua#ought>
"http""user""example.org"#f"/some/where/place""qua""ought"
All components except path filled out: <http://user@example.org:80?qua#ought>
"http""user""example.org"80#f"qua""ought"
All components except query filled out: <http://user@example.org:80/some/where/place#ought>
"http""user""example.org"80"/some/where/place"#f"ought"
All components except fragment filled out: <http://user@example.org:80/some/where/place?qua>
"http""user""example.org"80"/some/where/place""qua"#f
Empty host, nil user/port: <http:///some/where/place?qua#ought>
"http"#f""#f"/some/where/place""qua""ought"
Empty user, nil host/port: <http://@/some/where/place?qua#ought>
"http"""#f#f"/some/where/place""qua""ought"
Empty port implies empty host: <http://:/some/where/place?qua#ought>
"http"#f""#f"/some/where/place""qua""ought"
Relative reference, nil host: <////g>
N/A#f""#f"//g"#f#f
Relative reference, path, nil host: <.///g>
N/A#f#f#f".///g"#f#f
Relative reference, non-empty host: <//g>
N/A#f"g"#f#f#f#f
Path which looks like a query: <./p=q:r>
N/A#f#f#f"./p=q:r"#f#f
Relative reference, all components filled out: <//user@example.org:80/some/where/place?qua#ought>
N/A"user""example.org"80"/some/where/place""qua""ought"
Relative reference, all components except user filled out: <//example.org:80/some/where/place?qua#ought>
N/A#f"example.org"80"/some/where/place""qua""ought"
Relative reference, all components except host filled out: <//user@:80/some/where/place?qua#ought>
N/A"user"#f80"/some/where/place""qua""ought"
Relative reference, all components except port filled out: <//user@example.org/some/where/place?qua#ought>
N/A"user""example.org"#f"/some/where/place""qua""ought"
Relative reference, all components except path filled out: <//user@example.org:80?qua#ought>
N/A"user""example.org"80#f"qua""ought"
Relative reference, all components except query filled out: <//user@example.org:80/some/where/place#ought>
N/A"user""example.org"80"/some/where/place"#f"ought"
Relative reference, all components except fragment filled out: <//user@example.org:80/some/where/place?qua>
N/A"user""example.org"80"/some/where/place""qua"#f
Relative reference empty host, nil user/port: <///some/where/place?qua#ought>
N/A#f""#f"/some/where/place""qua""ought"
Relative reference empty user, nil host/port: <//@/some/where/place?qua#ought>
N/A""#f#f"/some/where/place""qua""ought"
Relative reference empty port implies empty host: <//:/some/where/place?qua#ought>
N/A#f""#f"/some/where/place""qua""ought"

Implementation

The sample implementation targets Chez Scheme and is written in (mostly) portable R6RS. It imports various SRFIs from the Chez-SRFI grab-bag. At the time of writing, the only external dependency is the sample implementation of the draft SRFI 262 pattern matcher [11].

References

Acknowledgements

Thanks to Ivan Raikov and Peter Bex for suggestions on improvements, especially with respect to rejecting invalid UTF-8 sequences, and the relativization procedures.

HTML formatting is lifted from SRFI 276 by Peter McGoron.

© 2026 Duncan Guthrie.

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice (including the next paragraph) shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.


Editor: Arthur A. Gleckler