by Duncan Guthrie
This SRFI is currently in draft status. Here is an explanation of each status that a SRFI can hold. To provide input on this SRFI, please send email to srfi-275@nospamsrfi.schemers.org. To subscribe to the list, follow these instructions. You can access previous messages via the mailing list archive.
This SRFI proposes a programming interface for working with RFC 3986 universal resource identifiers (URIs), as well as RFC 3987’s generalisation to internationalised resource identifiers (IRIs). This document defines record types, normalization procedures, and conversion between URIs and IRIs. The proposal also specifies a basic programming interface for working with paths in isolation, as this a pervasive usage of URIs and IRIs. Finally, we contribute a test suite to specify the behaviour of normalization with respect to relative references, which has in the past been a source of divergence between implementations.
none so far
RFC 3986 [1] describes an abstract syntax for uniform resource identifiers (URIs) as well as their relative references, which allow documents to be authored without knowing the final publishing location. RFC 3987 [2] defines IRIs, which generalise URIs to Unicode by defining an interpretation of percent-encoded bytes as sequences of UTF-8 octets.
URIs and IRIs are widely used to denote resources across the world wide web, and are the basis of a number of web standards. It is critical to present a programming interface for manipulation of different components in isolation (e.g. paths, hostnames), as working with an URI’s coarse string representation directly is error-prone. RFC 3986 defines an abstract syntax and hence conveniently forms the basis of such a programming interface. Further, RFC 3987’s generalisation to internationalised identifiers is a natural extension, allowing us to faithfully denote resources in a number of languages, using the universal character set. Indeed, IRIs form the basis of modern standards like Resource Description Framework (RDF), a widespread formal model for metadata interchange and knowledge representation. We hence require support for both URIs and IRIs.
Finally, paths are one of the most prevalent applications of URIs and IRIs. We develop paths as an object disjoint from other Scheme types with distinct segment structure, rather than adopting a string representation, because normalisation of paths is a key source of divergence in URI and IRI implementations to date. This object is located in a basic path library. This library structure also lets us expose procedures such as normalisation and merging of base and relative paths directly to the programmer.
RFC 3986 and 3987 distinguish URIs and IRIs, respectively, from their relative references, which must be resolved against a base URI or IRI in order to be used. The main application of relative references is to allow one to refer to resources and author documents without knowing the final publishing location. For example, a graph database may produce RDF documents specified in RDF/XML, where resources are denoted with IRI-relative references, assuming that these documents would be interchanged with another system which mints full IRIs with respect to its hosting location.
Somewhat confusingly, RFC 3986 defines "URI-reference" as the most common usage of URI: either an URI or a relative reference. We follow this common usage by developing polymorphic getters and setters which work on URIs and relative references, with the predicate uri? holding true for both URIs and relative references. This procedure would hence correspond to testing for an RFC 3986 "URI-reference". (RFC 3987 follows an identical convention for IRIs and their relative references, so the same approach is followed for IRIs.)
The most divergent behaviour between widespread implementations of URIs and IRIs has been with respect to normalization of relative references. We explicitly avoid defining path segment normalization for relative references (regardless of whether the path is absolute or relative) because the behaviour is largely undefined by RFC 3986 and 3987 or any current RFC. See the section on normalization for more details and the test suite.
URIs and IRIs are structured data. Of course, record types are convenient as they generate dedicated getters and setters. More importantly, however, we think that record structures are necessary in this case for normalization procedures to be structure-preserving. Specifically, normalization procedures may alter the path, such that its segments may be mistaken for other parts of the URI or IRI, such as the <://> portion separating scheme from authority, or for the hostname. Updates to authority components or to the path need to similarly ensure that they produce valid URIs and IRIs when serialised.
We argue that a string representation makes it too easy to inadvertently modify the structure, because normalization of individual components may yield a string which, when parsed again, is interpreted differently with respect to the structure. A good example of this is the restrictions on paths given an authority, because a non-empty authority is denoted in an URI using two slashes, which are characters which may also appear in paths.
A string representation of an URI or IRI (procedures uri->string and iri->string) is produced by concatenating string representations of the individual fields (scheme, hostname &c.), with the expected separators between components. For efficiency, no assumption should be made that the contents of individual fields can be checked at this point, which typically would involve additional, redundant parsing.
We require implementations to provide a purely-functional interface to URIs and IRIs, but not an in-place interface. The reasoning for this is that if implementors choose data structures optimised to purely-functional programming, it is more cumbersome to create an impure interface, whereas it is not as cumbersome for implementors to create a (potentially inefficient) purely-functional interface by copying the URI or IRI before updating in-place.
If an implementation provides disjoint mutable and immutable URI and IRI variants, then it is an error to call the in-place setters on an immutable variant, but the in-place variants should otherwise have equivalent error handling to the purely-functional variants.
Our design is to provide getters and setters which abstract away the internal representation, working on string representations of URI components. More generally, we suspect that existing implementations largely omit setters because preserving internal consistency is fairly cumbersome on the implementor and programmer, with validity of authority components being defined with respect to the path, and vice versa. The specific challenge is to ensure that setting a given component would not violate the URI grammar, as this may elicit a flat string representation which would have a different structure when parsed again.
In order to set one of the three authority components safely on an IRI or URI, then the following conditions must hold with respect to the IRI or URI’s existing path shape:
#f), rfc-path-abempty? must hold against the path.#f), and the IRI or URI is non-relative, then either rfc-path-absolute?, rfc-path-rootless? or rfc-path-empty? must hold against the path.#f), and the IRI or URI is a relative reference, then either rfc-path-absolute?, rfc-path-noscheme? or rfc-path-empty? must hold against the existing path component.Similarly, in order to set the path component safely, then the following conditions must hold with respect to the three authority components:
rfc-path-abempty? must hold against the path to be set.rfc-path-absolute?, rfc-path-rootless? or rfc-path-empty? must hold against the path to be set.rfc-path-absolute?, rfc-path-noscheme? or rfc-path-empty? must hold against the path to be set.Helper predicates for checking whether setting these components are located in the utility library: user-settable?, host-settable?, port-settable? and path-settable?.
We support three, scheme-independent normalization procedures:
.’ and ‘..’) are eliminated.The scheme and host components of both URIs and IRIs are considered case-insensitive, with other components considered case-sensitive. For URIs, the repertoire of characters is within U.S. ASCII., whereas for IRIs, Unicode characters may appear. Nonetheless, for both URIs and IRIs, only U.S. ASCII is case-insensitive, with characters like "É" never being normalized to a lower-case form like "é".
Escapes take the form of percent-encoded bytes, appearing as % 0-F 0-F (hexadecimal digits) in URIs and IRIs. These hexadecimal digits have a canonical upper-case form. For example, %cf should be normalized to %CF. In practice, the sample implementation always parses these into a canonical form as it represents these internally as octets, not in the original string form. If implementations do preserve the original form, they must always support normalization into the canonical upper-case form.
The interpretation of escapes (percent-encoded bytes) differs for URIs and IRIs.
For URIs, these octets may individually be decoded to ASCII characters. Essentially, this occurs if a character is not in the URI reserved range, and if it is permissible within a given URI component (e.g. path). Characters in the reserved range, if encountered in the clear, must never be escaped. Conversely, characters not in the reserved range, and which are not permissible within a given URI component, are encoded as a series of percent-encoded bytes corresponding to the character's UTF-8 octets. This bears particular mention because, while other character encodings are valid URIs, the RFC 3986 specification specifically requires this encoding, which enables compatibility with the closely related RFC 3987 specification for IRIs.
IRIs not only generalise the range of characters permissible within the IRI to certain Unicode ranges, but also interpret sequences of percent-encoded bytes as octets in UTF-8, which may be decoded into legal characters in the universal character set. Additionally, because all URIs are valid IRIs, the normalization of IRIs with respect to percent-encoded bytes is essentially the same as conversion from an URI to an IRI.
Path segment normalization is structurally identical for both URIs and IRIs. It interprets an entire path with respect to two control sequences: ‘.’ (current working directory) and ‘..’ (upper working directory), similar to UNIX paths. Unlike the other two normalization procedures, path segment normalization is undefined for relative references, because the path portion is not meaningful except during relative reference resolution.
It should be noted that path segment normalization is defined for non-relative IRIs with relative paths, such as a URN like <foo:a/b/../.././../../e>. This is a major source of diversion between RFC 3986 implementations. This appears to arise from reusing the Remove Dot Segments procedure as defined in RFC 3986 verbatim. In relative reference resolution, this procedure is never called on relative paths, only on absolute paths.
For example, for the (non-relative) IRI <foo:a/b/../.././../../e>, implementations using the RFC 3986 procedure verbatim get <foo:/e>, whereas other implementations get <foo:e>. In the former camp are implementations like the Erlang/OTP system’s built-in uri_string [3], and Guile-RDF [4], whereas in the latter camp are implementations like Haskell’s network-uri [5] and Chicken’s uri-generic [6] and uri-common [7] and libraries. We are in the latter camp, and a fixed Remove Dot Segments procedure might be implemented as follows. The sample implementation internally represents raw segments as a vector of SRFI 160 [10] u16vector segments, and uses the SRFI 262 pattern matcher [11].
(define (remove-dot-segments path)
(let* ([undotted
(remove-dot-segments/list (path-raw-segments path))]
[segments
(list->vector undotted)])
(cond [(absolute-path? path)
(make-raw-absolute-path segments)] ;; implementation-specific
[(and (relative-path? path)
(fx=? 1 (vector-length segments))
(u16vector-empty? (vector-ref segments 0)))
;; relative path with single empty segment is not meaningful:
(make-raw-relative-path (vector))] ;; implementation-specific
[else
(make-raw-relative-path segments)]))) ;; ''
(define (current-wd? seq) (u16vector= seq (u16vector #x2E)))
(define (parent-wd? seq) (u16vector= seq (u16vector #x2E #x2E)))
(define (remove-dot-segments/list segments)
(let loop ([ps (vector->list segments)]
[trailing-slash? #f]
[lst '()])
(match ps
['()
(if trailing-slash?
(reverse (cons (u16vector) lst))
(reverse lst))]
[(cons (? current-wd?) rst)
(loop rst #t lst)]
[(cons (? parent-wd?) rst)
(loop rst #t (if (pair? lst) (cdr lst) lst))]
[(cons x rst)
(loop rst #f (cons x lst))])))
Additionally, the above suggested implementation of remove-dot-segments is considerably clearer than the stack-based description in RFC 3986, with fewer pattern-matching clauses required.
Relative reference resolution resolves an URI or IRI against a (non-relative) base URI or IRI. This SRFI proposal simply names these procedures resolve-iri and resolve-uri as the "relative reference resolution" language in RFC 3986 and 3987 specifically refers to the URI-reference and IRI-reference (see above). Additionally, this naming is consistent with existing libraries such as Erlang's uri_string module [3] and Java's URI module [12].
The inverse of this procedure takes a non-relative URI or IRI and a base URI or IRI, and produces a relative reference which, given the base, would produce that non-relative URI or IRI. This is not defined in RFC 3986 or 3987, but is fairly widespread, such as relative-from found in Haskell's network-uri [5] and Chicken's uri-generic [6], and relativize found in Java's URI module [12]. The Haskell and Chicken libraries arguments are flipped for consistency with their direction, whereas Java's class methods belong to the base URI. This SRFI proposal names these procedures relativize-iri and relativize-uri, which follows Java's naming and argument convention, although the behaviour and test cases are based on the Chicken library:
Relativization proceeds as follows:
#f), and with the query and fragment of the non-base argument.path-rfc-abempty? holds. A relative reference is returned in which the authority components are unset, query is unset, and with the following path:
.’)Calculating that relative path, on the level of segments behaves as follows, or an equivalent procedure:
..’ segments. This sequence is our candidate relativized path segments..’).’ or ‘..’):
Finally, these segments are wrapped by a path object. The initial segment may be empty, in which case, the path object is absolute with segments excluding the initial segment. Otherwise, the path is a relative path using the segments calculated above. Additionally, the relativization procedures check that the path was settable given the authority component, and will prepend ./ to paths with a leading double slash given no authority component being set (see here).
We additionally specify a basic path sub-library. Paths can be considered vectors of segments which are tagged with whether they have a leading slash (an absolute path). The character set allowed within these paths is equivalent to the characters permissible within an URI or within the range of the universal character set (UCS) permitted within an IRI path segment. This restriction is important because it excludes a number of control characters and characters which can never be typed.
The basic path library specifies the basic operations of path creation from sets of segments (build-path), subscripting (path-ref), functional updates (path-update), and count of segments (path-length). We go a little further and provide utility procedures for path shape (the variants in the RFC 3986 and 3987 ABNF), and for copying paths. Finally, this design allows us to explicitly expose in the programming interface generalist procedures inherited from RFC 3986, such as merge-paths and remove-dot-segments.
For each procedure, error-handling behaviour is described in terms of exceptions to be raised and their irritants, either as an assertion violation, or as an error. R6RS conditions may not be available, in which implementations may simply signal an error. However, if these are available, then implementations should raise a compound condition as follows:
&who condition with the procedure name defined in this document&irritants condition with the irritants in the order described&assertion-violation or &error condition respectively, depending on the language in the procedure descriptionWith access to R6RS conditions, one might implement the encode-string procedure as follows. The first call to the R6RS assertion-violation procedure raises a compound condition of &who, &irritants, &assertion-violation and &error (as well as &message), where the who portion is the symbol encode-string, and the single irritant is the string argument.
(define encode-string
(case-lambda
[(str)
(encode-string str (char-set-complement char-set:iri-unreserved))]
[(str cset)
(cond [(not (string? str))
(assertion-violation 'encode-string "not a string" nonstr)]
[(not (char-set? escape-these))
(assertion-violation 'encode-string "not a character set" noncset)]
[else
...])]))
An R7RS-small implementation might implement it instead as follows:
(define encode-string
(case-lambda
[(str)
(encode-string str (char-set-complement char-set:iri-unreserved))]
[(str cset)
(cond [(not (string? str))
(error "not a string" nonstr)]
[(not (char-set? escape-these))
(error "not a character set" noncset)]
[else
...])]))
In RFC 3986 and 3987, IP literals (namely IPv6 addresses) appear in the host portion of an URI or IRI enclosed by square brackets. This SRFI proposal does not normalise these internally to an IP literal object or similar, with the IP literal returned by getters for the host portion enclosed by square brackets. This has the advantage that getters and setters exchange the same (bracketed) IP literal representation, receiving applications likely need to remove the brackets before usage of IP literals returned by this library's setters.
This specification follows the R6RS procedure entries convention. In addition to the naming conventions specifying type restriction for arguments where they are used, we add the following:
| iri | non-relative IRI or IRI relative reference |
| non-relative-iri | non-relative IRI |
| relative-iri | IRI relative reference |
| uri | non-relative URI or URI relative reference |
| non-relative-uri | non-relative URI |
| relative-uri | URI relative reference |
| path | relative or absolute path object |
| relative-path | relative path object |
| absolute-path | absolute path object |
Although we support both URIs and IRIs, for brevity we primarily describe behaviour for IRIs, and omit descriptions of the equivalent procedures for URIs where they behave the same. This works because URIs and IRIs are structurally identical, with the divergence between RFC 3986 and 3987 arising from the generalisation of the character set, and the treatment of normalization.
Library references are in the form (srfi :275 <sub-library>) consistent with SRFI 97 [9].
Finally, throughout this document, in examples we enclose IRIs and URIs in angle brackets, e.g. <http://example.org>. An object is assumed to be either an IRI or URI depending on the producing procedures, e.g. string->iri.
string->iri
string->non-relative-iri
string->rfc-iri-reference
string->rfc-absolute-iristring->uri
string->non-relative-uri
string->rfc-uri-reference
string->rfc-absolute-uriget-iri
get-non-relative-iri
get-rfc-iri-reference
get-rfc-absolute-iriget-uri
get-non-relative-uri
get-rfc-uri-reference
get-rfc-absolute-uriempty-iri empty-uriiri?
relative-iri?
non-relative-iri?
rfc-iri-reference?
rfc-absolute-iri?uri?
relative-uri?
non-relative-uri?
rfc-uri-reference?
rfc-absolute-uri?iri-equal? iri-eqv?uri-equal? uri-eqv?string->iri
iri->string
string->uri
uri->stringencode-string decode-stringiri->uri uri->iriiri-schemeiri-user iri-host iri-port iri-authority iri-username+passwordiri-path iri-path-stringiri-query iri-fragmenturi-schemeuri-user uri-host uri-port uri-authority uri-username+passworduri-path uri-path-stringuri-query uri-fragmentupdate-iri-schemeupdate-iri-user update-iri-host update-iri-port update-iri-authorityupdate-iri-path update-iri-query update-iri-fragmentupdate-uri-schemeupdate-uri-user update-uri-host update-uri-port update-uri-authorityupdate-uri-path update-uri-query update-uri-fragmentresolve-iri relativize-iriresolve-uri relativize-urinormalize-iri-case
normalize-iri-escape
normalize-iri-pathnormalize-uri-case
normalize-uri-escape
normalize-uri-pathbuild-path
vector->relative-path
vector->absolute-pathstring->path
string->relative-path
string->absolute-pathempty-relative-path
empty-absolute-pathvector->rfc-path-rootless
vector->rfc-path-noscheme
vector->rfc-path-abempty
vector->rfc-path-absolutestring->rfc-path-rootless
string->rfc-path-noscheme
string->rfc-path-abempty
string->rfc-path-absolutepath->string path-segmentspath?
path-string?
relative-path?
absolute-path?
empty-path?iri-path-relative?
iri-path-absolute?
iri-path-empty? uri-path-relative?
uri-path-absolute?
uri-path-empty? rfc-path-rootless?
rfc-path-noscheme?
rfc-path-abempty?
rfc-path-absolute?
rfc-path-empty?iri-path-rfc-rootless?
iri-path-rfc-noscheme?
iri-path-rfc-abempty?
iri-path-rfc-absolute?
iri-path-rfc-empty?uri-path-rfc-rootless?
uri-path-rfc-noscheme?
uri-path-rfc-abempty?
uri-path-rfc-absolute?
uri-path-rfc-empty?path-length
path-ref
path-update
path-equal?
remove-dot-segments
merge-pathsiri-path-segment-ref
iri-path-segment-update
iri-path-segmentsuri-path-segment-ref
uri-path-segment-update
uri-path-segmentsutf8->string/raise
percent-encoding->string
percent-encoding->u8-list
percent-encoding->reverse-u8-listusername+passworduser-settable?
host-settable?
port-settable?
path-settable?char-set:iri-unreserved
char-set:ucschar
char-set:iri-private
char-set:uri-unreservedchar-set:gen-delims
char-set:sub-delims char-set:reservedchar-set:schemechar-set:iri-userinfo char-set:uri-userinfochar-set:iri-reg-name char-set:uri-reg-namechar-set:iri-segment char-set:uri-segmentchar-set:iri-query char-set:uri-querychar-set:iri-fragment char-set:uri-fragmentchar-set:iri-allowed
char-set:uri-allowed
char-set:path-allowedReturns #t if obj is an IRI. Returns #f otherwise.
Returns #t providing that obj is a non-relative IRI (has a scheme component). Returns #f otherwise.
Returns #t providing that obj is an IRI relative reference (no scheme component). Returns #f otherwise.
Examples:
(define example-IRI (string->iri "/ex#IRI")) example-IRI → </ex#IRI> (iri? example-IRI) → #t (non-relative-iri? example-IRI) → #f (relative-iri? example-IRI) → #t
Exactly iri?, which already corresponds to RFC 3987’s IRI-reference production.
Returns #t providing that obj is a non-relative IRI and that its fragment component is #f. This corresponds exactly to RFC 3987’s absolute-IRI production. Returns #f otherwise.
Examples:
(define example0 (string->iri "http://example.org/ex?cond")) (define example1 (string->iri "http://example.org/ex#title")) (define example2 (string->iri "//example.org/ex?cond")) (define example3 (string->iri "//example.org/ex#title")) (map rfc-absolute-iri? (list example0 example1 example2 example3)) → (#t #f #f #f)
Helper procedure which takes no arguments and constructs a relative reference with all components set to #f, i.e. <> or the result of parsing "".
Retrieve the scheme component of non-relative-iri as a string. Unlike the procedures to follow, this procedure never returns #f as this would imply that the scheme is unset, i.e. a relative reference.
Retrieve the respective RFC 3987 fields of iri as either a string, or #f. An empty field is distinct from an unset one, for example given an IRI with bare ?, iri-query would return "", whereas if ? had been omitted, iri-query would return #f.
Retrieve the RFC 3987 port field of iri as either a non-negative integer, or #f.
Retrieve the RFC 3987 path field of iri as a path object as described in the path sub library. Unlike the other fields, #f is never returned.
Examples for the above seven component-specific procedures:
(define example-IRI (string->iri "http://example.org:80/ex#IRI"))
example-IRI → <http://example.org:80/ex#IRI>
(iri? example-IRI) → #t
(non-relative-iri? example-IRI) → #t
(iri-scheme example-IRI) → "http"
(iri-user example-IRI) → #f
(iri-host example-IRI) → "example.org"
(iri-port example-IRI) → 80
(path-segments (iri-path example-IRI)) → #("ex")
(iri-query example-IRI) → #f
(iri-fragment example-IRI) → "IRI"
Get the authority components of iri (user, host and port). This procedure returns #f if none of the three authority components are set, else all three as multiple values.
Helper procedure which splits the user field of iri at the first colon, producing strings corresponding to username and password as two values. If there is no colon then the whole user field is returned as first value, and #f as second. If the user field is not set, then #f and #f are returned. A colon appearing as a percent-encoded byte (%3A) is not considered a delimiter between username and password.
Examples:
(iri-username+password (string->iri"//foo:bar:qux@host")) → "foo" "bar:qux" (iri-username+password (string->iri"//foo%3Abar:qux@host")) → "foo%3Abar" "qux" (iri-username+password (string->iri"//@host")) → "" #f (iri-username+password (string->iri "//")) → #f #f (iri-username+password (string->iri "//foo@host")) → "foo" #f
We also provide a generic helper procedure for processing strings retrieved from IRIs or URIs user fields.
In contrast to iri-path, return a string representation of the path of iri, which may be the empty string.
Examples:
(define example0 (string->iri "http://example.org:80/ex#IRI"))
=(define example1 (string->iri "http://example.org:80#IRI"))
(define example2 (string->iri "//a"))
(path-segments (iri-path example0)) → #("ex")
(iri-path-string example0) → "/ex"
(path-segments (iri-path example1)) → #()
(iri-path-string example1) → ""
(path-segments (iri-path example2)) → #()
(iri-path-string example2) → ""
Purely-functional setter for the scheme component of non-relative IRI non-relative-iri, encoding the string string. During parsing, an error is raised if a character cannot be contained within a scheme at that position, with the character, its position in the IRI, and the IRI as irritants. Unlike the other setters for IRI components to follow, setting the scheme to #f or the empty string is not possible becasue it would imply that the IRI is a relative reference.
(update-iri-user iri obj)(update-iri-host iri obj)(update-iri-port iri obj)Purely-functional setters to set a respective authority components of IRI iri to obj. These procedures check that the authority component after being set to obj is not in conflict with the shape of the existing path. For instance, relative paths cannot usually be set when either of the authority components are set. The conditions associated with setting any authority component with the two input arguments are tested, and if they do not hold, then an error is raised with the input arguments as irritants.
In update-iri-user and update-iri-host, if obj is not #f and is a string, then it will be parsed for the respective component. Any character which cannot be contained within the component will be encoded as a sequence of percent-encoded bytes. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised. In update-iri-port, if obj is not #f and is a non-negative integer, then it will be set-as is, otherwise #f. An assertion violation is raised with obj as irritant if it is not of the expected type just discussed or #f.
Combination procedure which subsumes the above procedures. If a single argument #f is given, then all three components will be set to #f providing this is not in conflict with path. If three arguments are given, then error behaviour is identical to the component-specific setters except that the check for conflict with the existing path is that setting all three to would not be in conflict with the path shape.
Purely-functional setters for the query and fragment components of an IRI. Error behaviour is identical to update-uri-user, except that there is no check that these are in conflict with the path shape.
Update the path component of IRI iri with obj. If obj is #f, then the path to be set is the empty path. If obj is a path object, then any character not permissible within a valid IRI path is escaped as percent-encoded bytes corresponding to the charcter's UTF-8 octets. If obj is a string, then it is converted to a path object using a procedure equivalent to string->path, with any illegal characters again escaped as percent-encoded bytes. Next, the conditions associated with setting path are tested for iri and the path object to be set, and if they do not hold, then an error is raised with both input arguments as irritants. An assertion violation is raised with obj as irritant if it is not a path object, a string, or #f.
Returns #t if the iri1 and iri2 have fields which are equal. A non-relative IRI is not equal to a relative reference and vice versa as relative references have no scheme component. For components other than path and port, two components are only equal if they encode the same sequence of characters and percent-encoded bytes exactly. Two port components are equal if eq? holds between them, and two path componets are equal if path-equal? holds. Returns #f otherwise.
(uri? uri)(non-relative-uri? uri)(relative-uri? uri)(rfc-uri-reference? uri)(rfc-absolute-uri? uri)(empty-uri)(uri-scheme uri)(uri-user uri)(uri-host uri)(uri-query uri)(uri-fragment uri)(uri-port uri)(uri-path uri)(uri-authority uri)(uri-username+password uri)(uri-path-string uri)(update-uri-scheme uri obj)(update-uri-user uri obj)(update-uri-host uri obj)(update-uri-port uri obj)(update-uri-path uri obj)(update-uri-query uri obj)(update-uri-fragment uri obj)(update-uri-authority uri #f)(update-uri-authority uri user host port)(uri-equal? uri1 uri2)Identical programming interfaces to that of IRIs are given for URIs, with argument restriction to the respective URI types, and with the repertoire of characters escaped by setters extended to characters permissible in a given IRI component, but not the equivalent URI component.
Relative reference resolution of iri against base IRI base. This procedure is purely-functional as it (usually) involves transforming a relative reference into an IRI. An assertion violation is raised with irritant base if it is not a non-relative IRI. Simple examples derived from the RDF Turtle test cases (see test suite section):
(define string-cases (list "g:h" "g" "./g" "g/" "/g" "//g"))
(define base01 (string->iri "http://a/bb/ccc/d;p?q"))
(define base02 (string->iri "http://a/bb/ccc/d/"))
(define base07 (string->iri "file:///a/bb/ccc/d;p?q"))
(define (resolve-with base-iri) ;; Higher-order. Returns a function.
(lambda (ref)
(resolve-iri-reference base-iri (string->iri ref))))
(map (resolve-with base01) string-cases)
→ (list <g:h>
<http://a/bb/ccc/g> <http://a/bb/ccc/g> <http://a/bb/ccc/g/>
<http://a/g> <http://g>)
(map (resolve-with base02) string-cases)
→ (list <g:h>
<http://a/bb/ccc/d/g> <http://a/bb/ccc/d/g> <http://a/bb/ccc/d/g/>
<http://a/g> <http://g>)
(map (resolve-with base07) string-cases)
→ (list <g:h>
<file:///a/bb/ccc/g> <file:///a/bb/ccc/g> <file:///a/bb/ccc/g/>
<file:///g> <file://g>)
The inverse of resolve-iri, which elicits a relative reference which is resolveable against base in order to produce non-relative-iri. It is not possible to elicit a path for which rfc-path-abempty? does not hold after calling remove-dot-segments for either argument, and an error is raised with the normalized path and respective argument as irritants. An assertion violation is raised if base is not a non-relative IRI. Note that the argument order is flipped compared to Chicken uri-generic's [6] relative-from procedure.
Examples:
(define base-iri (string->iri "http://a/bb/ccc/d;p?q"))
(define reference0 (string->iri "g"))
(define resolved0 (string->iri "http://a/bb/ccc/g"))
;; (iri-equal? (relativize base (resolve base v)) v)
(iri-equal? (relativize-iri base-iri (resolve-iri base-iri reference0))
reference0) → #t
;; (iri-equal? (resolve base (relativize base v)) v)
(iri-equal? (resolve-iri base-iri (relativize-iri base-iri resolved0))
resolved0) → #t
Relative reference resolution of an IRI against a base URI, and the inverse, which produces a relative reference given a base IRI and a resolved IRI. Behaviour is structurally identical to that of relative reference resolution for IRIs, and the arguments are instead restricted to non-relative URIs and relative references respectively.
The following procedures are purely-functional and structure preserving on the level of components, to avoid some of the problems with string representations being parsed differently post-normalization.
Normalize case-insensitive components (scheme and host), and convert percent-encoded bytes to canonical upper-case.
Normalize an IRI’s escapes, potentially interpreting percent-encoded bytes if part of valid UTF-8 octet sequences as characters. Conversely, characters which are not permissible within an URI component will appear as series of percent-encoded bytes corresponding to UTF-8 octets. This procedure is idempotent: a fully normalized IRI will be normalized to itself, and iri-equal? will hold true between the two.
Normalize an URI’s escapes, potentially interpreting percent-encoded bytes as octets in U.S. ASCII. Conversely, characters which are not permissible within an URI component will appear as series of percent-encoded bytes corresponding to UTF-8 octets.
These procedures normalize path segments given the control segments ‘.’ and ‘..’. While path segments of IRI or URI-relative references are not normalized, these procedures simply have no effect and no error is signalled, consistent with the other normalization procedures. Although it is relatively uncommon for relative paths to appear in non-relative IRIs or URIs, they are permissible e.g. within URN components.
These procedures are essentially a sequence of the three normalization procedures described previously.
Process string, escaping any characters in the optional character set char-set, which defaults to characters outside of the IRI range (complement of char-set:iri-unreserved). Notably, percent-encoded bytes must not be interpreted as escapes, with the percent-sign treated as a character, and always percent-encoded in turn (as %25).
Process string, decoding any sequences of percent-encoded bytes as the corresponding UTF-8 octets for characters in the optional character set char-set, which defaults to every character (SRFI 14 char-set:full). When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error is raised with the #\% character and the start of the sequence in which the invalid sub-sequence occured as irritants.
(string->iri string)(string->non-relative-iri string)(string->rfc-iri-reference string)(string->rfc-absolute-iri string)string->iri parses string into an IRI. An error is raised with the failed character and its position in the string, in the following cases. First, this may occur when a character is not permitted within a given IRI component (unlike the string setters, there is no automatic escaping). Second, when a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised. Finally, if the end of the string was prematurely encountered, then the irritants are the EOF object and the string length.
string->non-relative-iri is a restriction on string->iri in that the scheme component is required. An error is raised with the EOF object and the string length as irritants if the scheme does not end with a colon, and with an illegal character and its position as irritants if this is contained within the scheme. Once a valid scheme has been parsed, the parse continues with the same behaviour as string->iri. This definition of a non-relative IRI (requirement of a scheme component) differs from the RFC 3987 absolute-IRI definition in that a fragment may appear, but we offer this procedure for convenience as many applications (like the RDF N-Triples concrete syntax) forbid relative references but permit non-relative IRIs to have fragment components.
Definitions derived from RFC 3987 verbatim are string->rfc-iri-reference and string->rfc-absolute-iri. The former procedure behaves exactly like string->iri but the naming reflects RFC 3987's IRI-reference terminology. The latter procedure has the same error behaviour as string->non-relative-iri except that the fragment component is strictly forbidden. That is, if the hash character is parsed, then an error is raised with the hash character and its position in the string as irritants.
Serialisation of an IRI as a string. Examples:
(define example-A (string->iri "http://example.org/some/where/place")) (define example-B (string->iri "urn:/some/where/place")) (iri->string example-A) → "http://example.org/some/where/place" (iri->string example-B) → "urn:/some/where/place"
(get-iri port)(get-iri port obj)(get-non-relative-iri port)(get-non-relative-iri port obj)(get-rfc-iri-reference port)(get-rfc-iri-reference port obj)(get-rfc-absolute-iri port)(get-rfc-absolute-iri port obj)These procedures are the equivalent of the string parsing procedures immediately above, but which parse a stream of text from a port as an IRI. These procedures stop if the port produces a character or the EOF object which is eq? with optional datum obj, which defaults to the EOF object. These low-level procedures are offered because they can be used directly in streaming parsers for data which contains IRIs, such as JSON-LD. All of these procedures raise an assertion violation if port is not a textual input port, and the parse behaviour is identical to the equivalent string procedures except that the position irritant is the equvialent port position.
(string->uri string)(string->non-relative-uri string)(string->rfc-uri-reference string)(string->rfc-absolute-uri string)(uri->string uri)(get-uri port)(get-uri port obj)(get-non-relative-uri port)(get-non-relative-uri port obj)(get-rfc-uri-reference port)(get-rfc-uri-reference port obj)(get-rfc-absolute-uri port)(get-rfc-absolute-uri port obj)Identical programming interface to that of IRIs are given for URIs, with argument restriction to the respective URI types, and with the repertoire of characters permissible unescaped restricted to those of URIs.
Two IRIs are equivalent if iri-equal? holds, or that it holds post-normalization with normalize-iri. Similarly, two URIs are equivalent if uri-equal? holds, or that it holds post-normalization with normalize-uri.
Convert an IRI to an URI. This proceeds by encoding any character within certain ranges (see RFC 3987ucschar and iprivate) to a series of percent-encoded bytes corresponding to the character's octets in UTF-8. This URI is also a valid IRI (albeit not normalized) as all URIs are valid IRIs. This procedure is structure-preserving: (non-relative) IRIs are never transformed into relative references, or vice-versa.
Convert an URI to an IRI. This procedure can be viewed as upgrading the URI structure to that of an IRI, then normalizing it as an IRI.
This section descibes the basic path sub-library, consisting of two path variants with a common internal segment structure, tagged with whether there is a leading slash. In addition to predicates to discinguish these two path types, we describe predicates which exactly follow the RFC 3986 and 3987 definitions, which are prefixed with rfc-. These do not check that the character set of the path is consistent with RFC 3986 or 3987, at most checking whether the first segment (if any) is empty or contains a colon character. Additionally, the rfc- procedures do not raise an error on non-path objects, instead returning #f, consistent with the generic absolute-path? or relative-path? procedures.
Predicates for relative and absolute path objects, which share a common internal segment structure, returning #t for the respective path objects. path? returns #t if its argument is a relative or absolute path. Returns #f otherwise.
Returns #t if a path is both relative and has no segments. An absolute path with no segments is not considered empty, as it starts with a slash and is represented by the string "/". Returns #f otherwise.
Returns #t if the input argument is a relative path with at least one, non-empty segment. Returns #f otherwise.
Returns #t if the input argument is a relative path with at least one, non-empty segment, with the additional condition that the initial segment does not contain a plain (unescaped) colon character. Returns #f otherwise.
Returns #t providing that either absolute-path? or empty-path?. Returns #f otherwise.
Returns #t providing the first segment is non-empty (i.e. the path does not start with //). Unlike rfc-path-abempty?, this procedure does not hold true for empty paths. Returns #f otherwise.
Wholly identical to empty-path?.
Examples for the above predicates:
(define rfc-abempty-example (string->path "//some/where/place"))
(define rfc-absolute-example (string->path "/some/where/place"))
(define rfc-rootless-example (string->path "some:thing/where/place"))
(define rfc-noscheme-example (string->path "some/where/place"))
(define relative-empty-example (string->path ""))
(define absolute-empty-example (string->path "/"))
(define all
(list rfc-abempty-example rfc-absolute-example
rfc-rootless-example rfc-noscheme-example
relative-empty-example absolute-empty-example))
(map absolute-path? all) → (#t #t #f #f #f #t)
(map relative-path? all) → (#f #f #t #t #t #f)
(map empty-path? all) → (#f #f #f #f #t #f)
(map rfc-path-abempty? all) → (#t #t #f #f #t #t) ; empty path is also rfc-path-abempty?
(map rfc-path-absolute? all) → (#f #t #f #f #f #f) ; zero segments is not rfc-path-absolute?
(map rfc-path-rootless? all) → (#f #f #t #t #f #f) ; empty path is not rfc-path-rootless?
(map rfc-path-noscheme? all) → (#f #f #f #t #f #f)
(map rfc-path-empty? all) → (#f #f #f #f #t #f)
Helper procedures which take no arguments and construct a relative path or an absolute path with no segments, i.e. a relative path like "" or an absolute path like "/"
Retrieve the length of the path segments, a non-negative integer.
Serialise the internal segment encoding as a vector of (string) segments. The returned vector cannot be guaranteed to correspond to the internal structure, as there may be percent-encoded bytes requiring a non-string encoding.
Retrieve segment of path path at index k as a string.
Update segment of path path at index k with string. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised.
Returns #t for two paths path1 and path2, providing that they are both absolute, or are both relative, and that compared pairwise, segments encode the exact same sequence of characters and percent-encoded bytes. Returns #f otherwise.
Copy the segment structure of a path object and tag it with whether path was relative or absolute.
Serialise path as a string, with no decoding of percent-encoded bytes.
Returns #t if x is either a path or string, and if a string, that it could represent a path without additional escaping. The specific check is that a character is permissible within an URI or is within the universal character set ranges permissible within an IRI (see char-set:path-allowed. No error is raised if x is not a string, instead #f is returned. This procedure corresponds to the similarly-named Racket procedure.
Parse string as a path. Both relative paths with a colon in the initial segment, and absolute paths with an empty initial segment, are permitted. When a percent-encoded byte is part of an invalid sequence of UTF-8 octets, an error consistent with utf8->string/raise is raised.
Convert string to a relative path. If the optional noscheme? argument is not #f, and the first segment contains a colon character, then the path output will be prefixed with ./. An error is raised with the string as irritant if it has a leading slash. Error behaviour is otherwise identical to string->path.
(string->absolute-path string)(string->absolute-path string nonempty-ini?)Convert a string to an absolute path. If the optional nonempt-ini? argument is not #f, and the initial segment is empty (leading double slash), then an error is raised with irritant string. Error behaviour is otherwise identical to string->path.
Helper procedures in which the final optional argument to string->relative-path is set to #f and #t respectively.
Helper procedures in which the final optional argument to string->absolute-path is set to #f and #t respectively.
Convert a vector of string segments to a relative path. This procedure behaves similarly to string->relative-path with the exception that that slashes are always escaped, so a slash at the start of the first segment does not raise an error. An error is raised if a segment is not a string, with the segment, its position, and vector as irritants, and if the initial segment is empty (serialised as leading slash), then an error is raised with the (string) segment, its position, and vector as irritants.
(vector->absolute-path vector)(vector->absolute-path vector nonempty-ini?)Convert a vector of string segments to an absolute path. This procedure behaves similarly to string->absolute-path with the exception that that slashes are always escaped, so a slash at the start of the first segment does not raise an error. An error is raised if a segment is not a string, with the segment, its position, and vector as irritants.
Helper procedures in which the final optional argument to vector->relative-path is set to #f and #t respectively.
Helper procedures in which the final optional argument to vector->absolute-path is set to #f and #t respectively.
Variadic procedure to build a path from a base path and any number of additional string segments. If no arguments are given, then the empty path is returned. This procedure corresponds to the similarly-named Racket procedure.
If base is a string, then the procedure behaves like vector->relative-path. If base is a path, then the procedure concatenates the base path with a relative path corresponding to the remaining segments. This is accomplished using the merge-paths procedure described in RFC 3986 section 5.2.3. The slash character is always escaped within segments.
By admitting the base as potentially another path, this allows one to build an absolute path from a list of segments, by choosing (empty-absolute-path) as base.
An error is raised if a segment is not a string, with the segment, and its position, as irritants. Additionally, if the initial segment (after expanding the base path) is empty, then an error is raised with the (string) segment and its position as irritants.
Remove dotted segments (‘.’ and ‘..’) in path by interpreting them alongside the path segment structure. This procedure corresponds to RFC 3986 section 5.2.4.
Merge base path base with path. As per RFC 3986 section 5.2.3, if base is empty and it has an authority component, then path is concatenated with the reference path. This is controlled by the whether the optional base-has-authority? argument is non-#f, and is mostly important during relative reference resolution. This defaults to #f as this procedure in isolation assumes no authority component.
(iri-path-absolute? iri)(iri-path-relative? iri)(iri-path-empty? iri)(iri-path-rfc-abempty? iri)(iri-path-rfc-absolute? iri)(iri-path-rfc-rootless? iri)(iri-path-rfc-noscheme? iri)(iri-path-rfc-empty? iri)Helper procedures which call the path library’s corresponding procedures for path shape of an IRI. Unlike the equvialent path procedures, these raise an assertion violation with iri as irritant if it is not an IRI. Examples:
(define rfc-abempty-example (string->iri "urn://myhost//some/where/place"))
(define rfc-absolute-example (string->iri "urn://myhost/some/where/place"))
(define rfc-rootless-example (string->iri "urn:some:thing/where/place"))
(define rfc-noscheme-example (string->iri "urn:some/where/place"))
(define relative-empty-example (string->iri "urn:"))
(define absolute-empty-example (string->iri "urn:/"))
(define all
(list rfc-abempty-example rfc-absolute-example
rfc-rootless-example rfc-noscheme-example
relative-empty-example absolute-empty-example))
(map iri-path-string all)
→ ("//some/where/place" "/some/where/place"
"some:thing/where/place" "some/where/place"
"" "/")
(map iri-path-absolute? all) → (#t #t #f #f #f #t)
(map iri-path-relative? all) → (#f #f #t #t #t #f)
(map iri-path-empty? all) → (#f #f #f #f #t #f)
(map iri-path-rfc-abempty? all) → (#t #t #f #f #t #t) ; empty path is also rfc-path-abempty?
(map iri-path-rfc-absolute? all) → (#f #t #f #f #f #f) ; zero segments is not rfc-path-absolute?
(map iri-path-rfc-rootless? all) → (#f #f #t #t #f #f) ; empty path is not rfc-path-rootless?
(map iri-path-rfc-noscheme? all) → (#f #f #f #t #f #f)
(map iri-path-rfc-empty? all) → (#f #f #f #f #t #f)
(iri-path-segments ident)(iri-path-segment-ref ident k)(iri-path-segment-update ident k str)Procedures which effectively wrap path-segments, path-ref and path-update respectively. These are offered to expose path internals without requiring to import the path sub-library. iri-path-segment-update has additional error behaviour beyond path-update: characters not permissible within an IRI path segment will be automatically escaped as sequences of percent-encoded bytes, comparable to the behaviour of update-iri-path on segments.
(uri-path-absolute? uri)(uri-path-relative? uri)(uri-path-empty? uri)(uri-path-rfc-abempty? uri)(uri-path-rfc-absolute? uri)(uri-path-rfc-rootless? uri)(uri-path-rfc-noscheme? uri)(uri-path-rfc-empty? uri)Helper procedures which call the path library’s corresponding procedures for path shape of an URI. Unlike the equvialent path procedures, these raise an assertion violation with ident as irritant if it is not an URI.
(uri-path-segments uri)(uri-path-segment-ref uri k)(uri-path-segment-update uri k str)Procedures which effectively wrap path-segments, path-ref and path-update respectively. In iri-path-segment-update, characters not permissible within an URI path segment will be automatically escaped.
R6RS’s utf8->string procedure silently substitutes the unicode replacement character (U+FFFD) given an invalid UTF-8 sequence. If the information is available, then this procedure raises an assertion violation with the offending octet and its position within the bytevector as irritants. Otherwise, the irritants are the offending octet alone.
Represent octet k (representing a percent-encoded byte) as a string in its canonical, upper-case form % 0-F 0-F. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256. Examples:
(percent-encoding->string #x40) → "%40" (percent-encoding->string #xFA) → "%FA" (percent-encoding->string #x00) → "%00" (percent-encoding->string -12.0 → ERROR) (percent-encoding->string #x100) → ERROR)
Represent octet k (reperesnting a percent-encoded byte) as a list of UTF-8 octets corresponding to the canonical string representation of the percent-encoded byte. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256. Examples:
(percent-encoding->u8-list #x40) → (#x25 #x34 #x30) (percent-encoding->u8-list #xFA) → (#x25 #x46 #x41) (percent-encoding->u8-list #x00) → (#x25 #x30 #x30) (map integer->char (percent-encoding->u8-list #xFA)) → (#\% #\F #\A)
Represent octet k (reperesnting a percent-encoded byte) as a list of UTF-8 octets corresponding to the canonical string representation of the percent-encoded byte, this time in reverse order. An assertion violation is raised with k as irritant if it is not an exact non-negative integer less than 256.
Split strings retrieved from an IRI or an URI’s user component into username and password components. Again, a colon which is escaped is not considered a delimiter between username and password. Examples:
(username+password (iri-user (string->iri"//foo:bar:qux@host"))) → "foo" "bar:qux" (username+password (iri-user (string->iri"//foo%3Abar:qux@host"))) → "foo%3Abar" "qux" (username+password (iri-user (string->iri"//@host"))) → "" #f (username+password (iri-user (string->iri "//"))) → #f #f (username+password (iri-user (string->iri "//foo@host"))) → "foo" #f
(user-settable? ident obj)(host-settable? ident obj)(port-settable? ident obj)(path-settable? ident obj)Helper procedures which hold providing that updating the respective component of IRI or URI ident to obj is legal. This directly reflects the RFC 3986 and 3987 ABNF, but it’s convenient to provide a single predicate as the conditions to check are somewhat elaborate with respect to path shape. These conditions are described in the error behaviour of setters section. For efficiency, obj is not parsed and the only check is that it is not #f. This means that subsequent procedures called on these arguments may error. An assertion violation is raised if the first argument is not an IRI or URI, with that argument as irritant.
Note that (srfi :275 iri) and (srfi :275 uri) export some of the same identifiers where the set of characters is the same in RFC 3986 and 3987: e.g. char-set:reserved. Character sets are named for the RFC 3986 and 3987 ABNF productions, e.g. char-set:uri-userinfo.
RFC 3986 unreserved character set, and the repertoire of all unescaped characters permissible in an URI.
RFC 3986 gen-delims and sub-delims character sets, and their union, the reserved range. These character sets are also exported by the (srfi :275 iri) library for convenience.
char-set:schemechar-set:uri-userinfochar-set:uri-reg-namechar-set:uri-segmentchar-set:uri-querychar-set:uri-fragmentRFC 3986 component-specific character sets. In RFC 3986, but not RFC 3987, query and fragment share the same repertoire.
char-set:iri-privatechar-set:ucscharchar-set:iri-unreservedchar-set:iri-allowedRFC 3987 iunreserved character set, which is the union of char-set:uri-unreserved and most of the universal character set. The permissible ranges in the UCS are exported as char-set:ucschar, and the additional char-set:iri-private range exports the additional characters permissible in RFC 3987 query components. Finally, char-set:iri-allowed is the repertoire of all characters permissible in an IRI.
char-set:schemechar-set:iri-userinfochar-set:iri-reg-namechar-set:iri-segmentchar-set:iri-querychar-set:iri-fragmentRFC 3987 component-specific character sets. The scheme component is the same as in RFC 3986. Query is more expansive in RFC 3987 and includes the char-set:iri-private range, unlike fragment.
Union of char-set:uri-allowed and char-set:ucschar described above. While this is minimally restritive, it does contain (a few) characters disallowed within IRI segments. However, any illegal characters will be escaped in the path setters for URIs and IRIs.
In this section, we describe various test cases and the specific behaviour they evaluate. Because RFC 3986 and 3987 only provide examples for relative reference resolution, it is important to specify the exact behaviour of a correct implementation, especially with respect to normalization. It is expected that these test cases could be the basis of a more comprehensive property-based test suite, e.g. test the entire range of reserved and unreserved characters in a particular URI component.
normalize-uri-case)<http://example.org/ex#test>
→ <http://example.org/ex#test><HttP://example.org/ex#test>
→ <http://example.org/ex#test><http://MySelf@example.org/Examp#test>
→ <http://MySelf@example.org/Examp#test><http://Example.ORG/ex#test>
→ <http://example.org/ex#test><http://example.org/Examp#test>
→ <http://example.org/Examp#test><http://example.org/examp?Qua#test>
→ <http://example.org/examp?Qua#test><http://example.org/examp#TeSt>
→ <http://example.org/examp#TeSt><http://%aA@%AA%AB%AC%AD%AE/some/where/place>
→ <http://%AA@%AA%AB%AC%AD%AE/some/where/place><http://%aa%Ab%AC%aD%AE/some/where/place>
→ <http://%AA%AB%AC%AD%AE/some/where/place><http://myname@example.org/%Fa/%FB/%fC>
→ <http://myname@example.org/%FA/%FB/%FC><http://myname@example.org/%FA/%FB/%FC?%ff>
→ <http://myname@example.org/%FA/%FB/%FC?%FF><http://myname@example.org/%FA/%FB/%FC#%ff>
→ <http://myname@example.org/%FA/%FB/%FC#%FF>normalize-iri-case)<http://CRÊPES.example.org>
→ <http://crÊpes.example.org>
No equivalent for scheme as that only contains ASCII even in IRIsnormalize-uri-escape)<http://my!name@example.org/ex#test>
→ <http://my!name@example.org/ex#test><http://myname@!example.org/ex#test>
→ <http://myname@!example.org/ex#test><http://myname@example.org/ex!#test>
→ <http://myname@example.org/ex!#test><http://myname@example.org/ex?!a#test>
→ <http://myname@example.org/ex?!a#test><http://myname@example.org/ex?a#!test>
→ <http://myname@example.org/ex?a#!test><http://my%40name@example.org/ex#test>
→ <http://my%40name@example.org/ex#test><http://myname@ex%40ample.org/ex#test>
→ <http://myname@ex%40ample.org/ex#test><http://myname@example.org/e%40x#test>
→ <http://myname@example.org/e%40x#test><http://myname@example.org/ex?a%40#test>
→ <http://myname@example.org/ex?a%40#test><http://myname@example.org/ex?a#t%40est>
→ <http://myname@example.org/ex?a#t%40est><http://my%2Ename@example.org/ex?a#test>
→ <http://my.name@example.org/ex?a#test><http://myname@example%2Eorg/ex?a#test>
→ <http://myname@example.org/ex?a#test><http://myname@example.org/misc%2Etxt#test>
→ <http://myname@example.org/misc.txt#test><http://myname@example.org/misc.txt?%2E%2E%2E>
→ <http://myname@example.org/misc.txt?...><http://myname@example.org/misc.txt#line%31%30>
→ <http://myname@example.org/misc.txt#line10><http://dosh£@crepes.example.org>
→ <http://dosh%C2%A3@crepes.example.org><http://crêpes.example.org>
→ <http://cr%C3%AApes.example.org><http://crepes.example.org/in/Rhône>
→ <http://crepes.example.org/in/Rh%C3%B4ne><http://crepes.example.org/in/Rennes?Dim.‥Sam.>
→ <http://crepes.example.org/in/Rennes?Dim.%E2%80%A5Sam.><http://crepes.example.org/in/Rennes#L'Étage>
→ <http://crepes.example.org/in/Rennes#L'%C3%89tage>normalize-iri-escape)<http://dosh%C2%A3@crepes.example.org>
→ <http://dosh£@crepes.example.org><http://cr%C3%AApes.example.org>
→ <http://crêpes.example.org><http://crepes.example.org/in/Rh%C3%B4ne>
→ <http://crepes.example.org/in/Rhône><http://crepes.example.org/in/Rennes?Dim.%E2%80%A5Sam.>
→ <http://crepes.example.org/in/Rennes?Dim.‥Sam.><http://crepes.example.org/in/Rennes#L'%C3%89tage>
→ <http://crepes.example.org/in/Rennes#L'Étage><https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82>
→ <https://en.wiktionary.org/wiki/Ῥόδος><https://example.org/music/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh>
→ <https://example.org/music/Éirigh'sCuirOrtDoChuidÉadaigh><https://en.wiktionary.org/wiki/Ῥόδος>
→ <https://en.wiktionary.org/wiki/Ῥόδος><https://en.wiktionary.org/wiki/Ῥόδος>
→ <https://en.wiktionary.org/wiki/Ῥόδος>
In this test case, normalize-iri-escape should be called multiple times, i.e. (compose normalize-iri-escape normalize-iri-escape)normalize-uri-path and normalize-iri-path)<http://example.org/some/where/place>
→ <http://example.org/some/where/place><urn:/some/where/place>
→ <urn:/some/where/place><urn:some/where/place>
→ <urn:some/where/place><urn:/some/./where/././place/./>
→ <urn:/some/where/place/><urn:some/./where/././place/./>
→ <urn:some/where/place/><urn:/some//where//place//>
→ <urn:/some//where//place//><urn:some//where//place//>
→ <urn:some//where//place//></>
→ </><//>
→ <//></a/b/../../c>
→ </a/b/../../c></a/b/././c>
→ </a/b/././c></a/b/../c/././d>
→ </a/b/../c/././d><a/b/../../c>
→ <a/b/../../c><a/b/././c>
→ <a/b/././c><a/b/../c/././d>
→ <a/b/../c/././d><./def>
→ <./def><./abc:def>
→ <./abc:def><../../abc/./def>
→ <../../abc/./def><foo:a/b/../.././../../e>
→ <foo:e>
From Haskell network-uri [5]<http://example.com////../..>
→ <http://example.com//>
From Webkit [13]<http://example.com/foo/bar//../..>
→ <http://example.com/foo/>
From Webkit [13]<http://example.com/foo/bar//..>
→ <http://example.com/foo/bar/>
From Webkit [13]<http://example/a/b/../../c>
→ <http://example/c>
From Haskell network-uri [5]<http://example/a/b/c/../../>
→ <http://example/a/>
From Haskell network-uri [5]<http://example/a/b/c/./>
→ <http://example/a/b/c/>
From Haskell network-uri [5]<http://example/a/b/c/.././>
→ <http://example/a/b/>
From Haskell network-uri [5]<http://example/a/b/c/d/../../../../e>
→ <http://example/e>
From Haskell network-uri [5]<http://example/a/b/c/d/../.././../../e>
→ <http://example/e>
From Haskell network-uri [5]<http://example/a/b/../.././../../e>
→ <http://example/e>
From Haskell network-uri [5]uri->iri)<https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82>
→ <https://en.wiktionary.org/wiki/Ῥόδος><https://example.org/ceol/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh>
→ <https://example.org/ceol/Éirigh'sCuirOrtDoChuidÉadaigh>iri->uri)<https://en.wiktionary.org/wiki/Ῥόδος>
→ <https://en.wiktionary.org/wiki/%E1%BF%AC%CF%8C%CE%B4%CE%BF%CF%82><https://example.org/ceol/Éirigh'sCuirOrtDoChuidÉadaigh>
→ <https://example.org/ceol/%C3%89irigh'sCuirOrtDoChuid%C3%89adaigh>resolve-iri and resolve-uri)The following RDF 1.1 Turtle test cases for IRI resolution must succeed. These tests are based on those of RFC 3986, and are structured as triples in which the subject and predicate components are absolute IRIs (URNs), but where the object component is a relative IRI. A base IRI is declared at the top of the Turtle (.ttl) file, and the corresponding N-Triples (.nt) file lists the resolved triples. In the SRFI 275 sample implementation, we inline these tests, with the objects extracted and resolved against the base IRI explicitly.
<http://a/bb/ccc/d;p?q>)<http://a/bb/ccc/d/>)<file:///a/bb/ccc/d;p?q>)We require the following abnormal test cases from Chicken uri-generic [6], when resolved against base IRI <http://a/b/c/d;p?q>, to yield the following results:
..’ traversal is clamped at root, not an error<../../../g> → <http://a/g><../../../../g> → <http://a/g><../../../..> → <http://a/><../../../../> → <http://a/></./g> → <http://a/g></../g> → <http://a/g><g..> → <http://a/b/c/g..><..g> → <http://a/b/c/..g><g?y/./x> → <http://a/b/c/g?y/./x><g?y/../x> → <http://a/b/c/g?y/../x><g#s/./x> → <http://a/b/c/g#s/./x><g#s/../x> → <http://a/b/c/g#s/../x>Reference resolution must ensure that a path is never elicited which would be mistaken for the double slash after a scheme’s colon. (The selected strategy is to prefix a ./ to the path, making it relative, but in principle prepending a slash to the segments i.e. after the initial slash, /.///, is equivalent.)
<f:/a><.//g> → <f:.///g><f:/a/><..//g> → <f:.///g>The following test cases are derived from the SWAP project's [8] uripath.py These tests are formatted here as base URI or IRI, relative reference, and expected result.
<foo:xyz><bar:abc> → <bar:abc><http://example/x/y/z><../abc> → <http://example/x/abc><http://example2/x/y/z><//example/x/abc> → <http://example/x/abc><http://ex/x/y/z><../r> → <http://ex/x/r><http://ex/x/y><q/r> → <http://ex/x/q/r><http://ex/x/y><q/r#s> → <http://ex/x/q/r#s><http://ex/x/y><q/r#s/t> → <http://ex/x/q/r#s/t><http://ex/x/y><ftp://ex/x/q/r> → <ftp://ex/x/q/r><http://ex/x/y><y> → <http://ex/x/y><http://ex/x/y/><.> → <http://ex/x/y/><http://ex/x/y/pdq><pdq> → <http://ex/x/y/pdq><http://ex/x/y/><z/> → <http://ex/x/y/z/><file:/swap/test/animal.rdf><animal.rdf#Animal> → <file:/swap/test/animal.rdf#Animal><file:/e/x/y/z><../abc> → <file:/e/x/abc><file:/example2/x/y/z><../../../example/x/abc> → <file:/example/x/abc><file:/ex/x/y/z><../r> → <file:/ex/x/r><file:/ex/x/y/z><../../../r> → <file:/r><file:/ex/x/y><q/r> → <file:/ex/x/q/r><file:/ex/x/y><q/r#s> → <file:/ex/x/q/r#s><file:/ex/x/y><q/r#> → <file:/ex/x/q/r#><file:/ex/x/y><q/r#s/t> → <file:/ex/x/q/r#s/t><file:/ex/x/y><ftp://ex/x/q/r> → <ftp://ex/x/q/r><file:/ex/x/y><y> → <file:/ex/x/y><file:/ex/x/y/><.> → <file:/ex/x/y/><file:/ex/x/y/pdq><pdq> → <file:/ex/x/y/pdq><file:/ex/x/y/><z/> → <file:/ex/x/y/z/><file:/devel/WWW/2000/10/swap/test/reluri-1.n3><//meetings.example.com/cal#m1> → <file://meetings.example.com/cal#m1><file:/home/connolly/w3ccvs/WWW/2000/10/swap/test/reluri-1.n3><//meetings.example.com/cal#m1> → <file://meetings.example.com/cal#m1><file:/some/dir/foo><.#blort> → <file:/some/dir/#blort><file:/some/dir/foo><.#> → <file:/some/dir/#><http://example/x/y%2Fz> (see here)<abc> → <http://example/x/abc><http://example/x/y/z> (see here)<../../x%2Fabc> → <http://example/x%2Fabc><http://example/x/y%2Fz> (see here)<../x%2Fabc> → <http://example/x%2Fabc><http://example/x%2Fy/z> (see here)<abc> → <http://example/x%2Fy/abc><http://example/x/abc.efg><.> → <http://example/x/>The following test cases are derived from the Haskell network-uri [5] package. These tests are formatted here as base URI or IRI, relative reference, and expected result. The first set of cases are tricky cases of non-relative paths together with query and fragment.
<mailto:local1@domain1?query1><local2@domain2> → <mailto:local2@domain2><mailto:local1@domain1><local2@domain2?query2> → <mailto:local2@domain2?query2><mailto:local1@domain1?query1><local2@domain2?query2> → <mailto:local2@domain2?query2><mailto:local@domain?query1><?query2> → <mailto:local@domain?query2><mailto:?query1><local@domain?query2> → <mailto:local@domain?query2><mailto:local@domain?query1><?query2> → <mailto:local@domain?query2><foo:bar><http://example/a/b?c/../d> → <http://example/a/b?c/../d><foo:bar><http://example/a/b#c/../d> → <http://example/a/b#c/../d>The remaining set of tests from network-uri [5] are with respect to dealing with the final segment when inverting (see next section). See here. These test cases are not exactly invertible, however as they do not all produce the exact same reference due to relative segments.
<http://www.example.com/data/limit/..><test.xml> → <http://www.example.com/data/limit/test.xml><file:/some/dir/foo><./#blort> → <file:/some/dir/#blort><file:/some/dir/foo><./#> → <file:/some/dir/#><file:/some/dir/..><./#blort> → <file:/some/dir/#blort><http://example.org/base/uri><http:this> → <http:this><http:base><http:this> → <http:this><f://example.org/base/a><b/c//d/e> → <f://example.org/base/b/c//d/e><mid:m@example.ord/c@example.org><m2@example.ord/c2@example.org> → <mid:m@example.ord/m2@example.ord/c2@example.org><file:///C:/DEV/Haskell/lib/HXmlToolbox-3.01/examples/><mini1.xml> → <file:///C:/DEV/Haskell/lib/HXmlToolbox-3.01/examples/mini1.xml><foo:a/y/z><../b/c> → <foo:a/b/c>relativize-iri and relativize-uri)These procedures are effectively the inverse of resolve-iri and resolve-uri. However, not all of the test cases described above succeed verbatim, as e.g. the references may include dotted segments which would be noramlised during the relative reference resolution process. Of the test cases descibed above, we specifically require that the test cases derived from the Python SWAP project as all of these test cases are invertible exactly.
Recall that test data above are formatted as a base URI or IRI, a relative reference, and an expected result of reference resolution. A test of relative reference extraction succeeds if given the base IRI or URI and the expected result of resolution, the relative reference (described in the middle column) is elicited.
The following inversion must hold, but note that we elicit relative reference </g>, which is only the same as previous reference <.//g> after calling remove-dot-segments.
<f:/a><f:.///g> → </g> <f:/a/><f:.///g> → <..//g>Finally, the following test cases included with Chicken's uri-generic library [6] must hold for base <http://a/b/c/d;p?q>:
<http://a/b/c> → <../c></>, but that's not convenient<http://a/> → <../..><http://a> → <//a><ftp://a/b/c/d;p?q> → <ftp://a/b/c/d;p?q><ftp://x/y/z;a?b> → <ftp://x/y/z;a?b><http://a/b/c/d;p?q> → <d;p><http://a/b/c/e> → <e><http://a/b/c/> → <.><http://a/b/e> → <../e><http://a/b/> → <..><http://b> → <//b><http://b/> → <//b/><http://b/c> → <//b/c>In this section, we detail negative test cases which must be rejected. These go beyond the RFC 3986 and 3987 specifications which permit percent-encoded bytes which may correspond to invalid UTF-8 octet sequences. No encodings other than UTF-8 are supported. The following example cases are for the URI component. The same sequences must be rejected in every component in which percent-encoded bytes may appear (every component except scheme and port).
(encode-string "a béc♂d😎e" (char-set-complement char-set:uri-allowed))
→ "a%20b%C3%A9c%E2%99%82d%F0%9F%98%8Ee"(decode-string "a%20b%C3%A9c%E2%99%82d%F0%9F%98%8Ee" (char-set-complement char-set:uri-allowed))
→ "a béc%E2%99%82d😎e"(encode-string "foo%20bar" (char-set-difference char-set:full char-set:uri-allowed))
→ "foo%2520bar"(decode-string "foo%2520bar" char-set:full)
→ "foo%20bar"The following strings must be rejected by decode-string (in contrast encode-string does not interpret escapes):
"a%C0%A0b""a%E0%80%A0b""a%E0%80%80%A0b""a%E0%80%A9b""a%C0b""a%E2%99""a%F0%9F%98""a%F0%9F%98x""a%A9"Additionally, we require that these sequences are rejected by the different parsers, as well as setters for IRI and URI components (e.g. update-iri-host). The utf8->string/raise procedure is provided for detecting these sequences in external code.
We expect the following edge-cases to have the corresponding hostname verbatim:
<http://[::]/> → "[::]"<http://[::1]/> → "[::1]"::<http://[1::]/"> → "[1::]"::, prefix<http://[2001:db8::]/> → "[2001:db8::]"::<http://[::2001:db8]/> → "[::2001:db8]"<http://[::ffff:192.0.2.1]/> → "[::ffff:192.0.2.1]"<http://[64:ff9b::192.0.2.1]/> → "[64:ff9b::192.0.2.1]"<http://[2001:db8::1]:8080/path?query#fragment> → "[2001:db8::1]"Negative test cases which must be rejected:
:: compression marker<http://[2001:db8:::1]/><http://[2001:db8:a:b:c:d:e:f:1]/>:: (too few)< ><http://[2001:db8::fffff]/><http://[2001:00db8::0001]/><http://[2001:db8::192.0.2]/>:: markers<http://[2001:db8:a::b::c]/>::<http://[2001:db8]/><http://[2001:db8>| scheme | user | host | port | path | query | fragment |
|---|---|---|---|---|---|---|
Empty URI: <> | ||||||
| N/A | #f | #f | #f | #f | #f | #f |
Empty authority: <//> | ||||||
| N/A | #f | "" | #f | #f | #f | #f |
Empty user: <//@> | ||||||
| N/A | "" | #f | #f | #f | #f | #f |
Empty port: <//:> | ||||||
| N/A | #f | "" | #f | #f | #f | #f |
Empty query: <?> | ||||||
| N/A | #f | #f | #f | #f | "" | #f |
Empty fragment: <#> | ||||||
| N/A | #f | #f | #f | #f | #f | "" |
Path which looks like a hostname: <example.org> | ||||||
| N/A | #f | #f | #f | "example.org" | #f | #f |
URN-like: <urn:something> | ||||||
"urn" | #f | #f | #f | "something" | #f | #f |
URN-like, path looks like hostname: <urn:example.org> | ||||||
"urn" | #f | #f | #f | "example.org" | #f | #f |
Path which looks like a URN: <./urn:something> | ||||||
| N/A | #f | #f | #f | "./urn:something" | #f | #f |
User with colon segment: <http://a:b@c:29> | ||||||
"http" | "a:b" | "c" | 29 | #f | #f | #f |
User-like component appears as path: <http::@c:29> | ||||||
"http" | #f | #f | #f | ":@c:29" | #f | #f |
Host-like component appears as user: <http://example.org:b@d/> | ||||||
"http" | "example.org:b" | "d" | #f | "/" | #f | #f |
Padded port as numeric value: <http://example.org:000080> | ||||||
"http" | #f | "example.org" | 80 | #f | #f | #f |
Query component with question mark: <http://example.org/abcd?efgh?ijkl> | ||||||
"http" | #f | "example.org" | #f | "/abcd" | "efgh?ijkl" | #f |
Fragment component with question mark: <http://example.org/abcd#efgh?ijkl> | ||||||
"http" | #f | "example.org" | #f | "/abcd" | #f | "efgh?ijkl" |
Path where first segment looks like host: <http:///some/where/place> | ||||||
"http" | #f | "" | #f | "/some/where/place" | #f | #f |
Scheme with nil host: <foo:> | ||||||
"foo" | #f | #f | #f | #f | #f | #f |
Scheme with path, empty host: <foo:////g> | ||||||
"foo" | #f | "" | #f | "//g" | #f | #f |
Scheme with path, nil host: <foo:.///g> | ||||||
"foo" | #f | #f | #f | ".///g" | #f | #f |
Scheme with non-empty host: <foo://g> | ||||||
"foo" | #f | "g" | #f | #f | #f | #f |
All components filled out: <http://user@example.org:80/some/where/place?qua#ought> | ||||||
"http" | "user" | "example.org" | 80 | "/some/where/place" | "qua" | "ought" |
All components except user filled out: <http://example.org:80/some/where/place?qua#ought> | ||||||
"http" | #f | "example.org" | 80 | "/some/where/place" | "qua" | "ought" |
All components except host filled out: <http://user@:80/some/where/place?qua#ought> | ||||||
"http" | "user" | #f | 80 | "/some/where/place" | "qua" | "ought" |
All components except port filled out: <http://user@example.org/some/where/place?qua#ought> | ||||||
"http" | "user" | "example.org" | #f | "/some/where/place" | "qua" | "ought" |
All components except path filled out: <http://user@example.org:80?qua#ought> | ||||||
"http" | "user" | "example.org" | 80 | #f | "qua" | "ought" |
All components except query filled out: <http://user@example.org:80/some/where/place#ought> | ||||||
"http" | "user" | "example.org" | 80 | "/some/where/place" | #f | "ought" |
All components except fragment filled out: <http://user@example.org:80/some/where/place?qua> | ||||||
"http" | "user" | "example.org" | 80 | "/some/where/place" | "qua" | #f |
Empty host, nil user/port: <http:///some/where/place?qua#ought> | ||||||
"http" | #f | "" | #f | "/some/where/place" | "qua" | "ought" |
Empty user, nil host/port: <http://@/some/where/place?qua#ought> | ||||||
"http" | "" | #f | #f | "/some/where/place" | "qua" | "ought" |
Empty port implies empty host: <http://:/some/where/place?qua#ought> | ||||||
"http" | #f | "" | #f | "/some/where/place" | "qua" | "ought" |
Relative reference, nil host: <////g> | ||||||
| N/A | #f | "" | #f | "//g" | #f | #f |
Relative reference, path, nil host: <.///g> | ||||||
| N/A | #f | #f | #f | ".///g" | #f | #f |
Relative reference, non-empty host: <//g> | ||||||
| N/A | #f | "g" | #f | #f | #f | #f |
Path which looks like a query: <./p=q:r> | ||||||
| N/A | #f | #f | #f | "./p=q:r" | #f | #f |
Relative reference, all components filled out: <//user@example.org:80/some/where/place?qua#ought> | ||||||
| N/A | "user" | "example.org" | 80 | "/some/where/place" | "qua" | "ought" |
Relative reference, all components except user filled out: <//example.org:80/some/where/place?qua#ought> | ||||||
| N/A | #f | "example.org" | 80 | "/some/where/place" | "qua" | "ought" |
Relative reference, all components except host filled out: <//user@:80/some/where/place?qua#ought> | ||||||
| N/A | "user" | #f | 80 | "/some/where/place" | "qua" | "ought" |
Relative reference, all components except port filled out: <//user@example.org/some/where/place?qua#ought> | ||||||
| N/A | "user" | "example.org" | #f | "/some/where/place" | "qua" | "ought" |
Relative reference, all components except path filled out: <//user@example.org:80?qua#ought> | ||||||
| N/A | "user" | "example.org" | 80 | #f | "qua" | "ought" |
Relative reference, all components except query filled out: <//user@example.org:80/some/where/place#ought> | ||||||
| N/A | "user" | "example.org" | 80 | "/some/where/place" | #f | "ought" |
Relative reference, all components except fragment filled out: <//user@example.org:80/some/where/place?qua> | ||||||
| N/A | "user" | "example.org" | 80 | "/some/where/place" | "qua" | #f |
Relative reference empty host, nil user/port: <///some/where/place?qua#ought> | ||||||
| N/A | #f | "" | #f | "/some/where/place" | "qua" | "ought" |
Relative reference empty user, nil host/port: <//@/some/where/place?qua#ought> | ||||||
| N/A | "" | #f | #f | "/some/where/place" | "qua" | "ought" |
Relative reference empty port implies empty host: <//:/some/where/place?qua#ought> | ||||||
| N/A | #f | "" | #f | "/some/where/place" | "qua" | "ought" |
The sample implementation targets Chez Scheme and is written in (mostly) portable R6RS. It imports various SRFIs from the Chez-SRFI grab-bag. At the time of writing, the only external dependency is the sample implementation of the draft SRFI 262 pattern matcher [11].
network-uri package.uri-genericuri-commonThanks to Ivan Raikov and Peter Bex for suggestions on improvements, especially with respect to rejecting invalid UTF-8 sequences, and the relativization procedures.
HTML formatting is lifted from SRFI 276 by Peter McGoron.
© 2026 Duncan Guthrie.
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice (including the next paragraph) shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.