From: Patrick Steinhardt <hidden> Date: 2017-01-23 13:13:09
Hi,
Short disclaimer: this patch results from work for a client at my
day job at elego Software Solutions GmbH. As such, I'm using my
work mail address and added a new mailmap entry. I wasn't exactly
certain if the mailmap entry should've been created in a separate
commit series, as it has nothing to do with the actual topic --
I can re-send it separately if requested.
This patch is mostly a request for comments. The use case is to
be able to configure an HTTP proxy for all subdomains of a
certain domain where there are hundreds of subdomains. The most
flexible way I could imagine was by using regular expressions for
the matching, which is how I implemented it for now. So users can
now create a configuration key like
`http.?http://.*\\.example\\.com.*` to apply settings for all
subdomains of `example.com`.
I tried to make this feature as backwards-compatible as it can be
by having the '?' prefix. Older clients will barf when trying to
normalize the URL as '?' is not in the set of allowed characters
for a URL, and for newer clients there will be no change in
behavior for previously configured `http.<url>.*` keys.
Regards
Patrick Steinhardt
Patrick Steinhardt (2):
mailmap: add Patrick Steinhardt's work address
urlmatch: allow regex-based URL matching
.mailmap | 1 +
Documentation/config.txt | 6 ++++-
t/t1300-repo-config.sh | 31 ++++++++++++++++++++++++++
urlmatch.c | 57 ++++++++++++++++++++++++++++++++++++++----------
4 files changed, 82 insertions(+), 13 deletions(-)
--
2.11.0
From: Patrick Steinhardt <hidden> Date: 2017-01-23 13:13:04
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to have regex-based URL matches. A
user can prefix a configuration key's URL with a question mark ('?') to
use regular expressions instead of exact matches in order to find all
matching URLs. A user can now simply add a key
`http.?http://.*\\.example\\.com.proxy` to set a proxy for all
subdomains of `example.com`. When no question mark is given as a prefix,
then the configuration subsystem will use the old algorithm based on
exact matches.
Signed-off-by: Patrick Steinhardt <redacted>
---
Documentation/config.txt | 6 ++++-
t/t1300-repo-config.sh | 31 ++++++++++++++++++++++++++
urlmatch.c | 57 ++++++++++++++++++++++++++++++++++++++----------
3 files changed, 81 insertions(+), 13 deletions(-)
@@ -1906,7 +1906,11 @@ http.followRedirects:: http.<url>.*:: Any of the http.* options above can be applied selectively to some URLs.- For a config key to match a URL, each element of the config key is+ There are two different modes to match URLs: if the config key's URL is+ prefixed with a `?`, it allows to make use of regular expressions. An+ example for this is `http.?http://.*\\.example\\.com.*` to match all+ subdomains of `example.com`.+ If the key is not prefixed with a `?`, each element of the config key is compared to that of the URL, in the following order: + --
@@ -1177,6 +1177,37 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'+test_expect_success'regex-based urlmatch''+cat>.git/config<<-\EOF&&+[http]+sslVerify+[http"?https://.*\\.example\\.com"]+sslVerify=false+cookieFile=/tmp/cookie.txt+EOF++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good.example.com>actual&&+test_must_be_emptyactual&&++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good-example.com>actual&&+test_must_be_emptyactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://example.com>actual&&+test_cmpexpectactual&&++echofalse>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://subdomain.example.com>actual&&+test_cmpexpectactual&&++{+echohttp.cookiefile/tmp/cookie.txt&&+echohttp.sslverifyfalse+}>expect&&+gitconfig--get-urlmatchHTTPhttps://subdomain.example.com>actual&&+test_cmpexpectactual+'+# good section hygiene test_expect_failure'unsetting the last key in a section removes header''cat>.git/config<<-\EOF&&
From: Patrick Steinhardt <hidden> Date: 2017-01-24 17:00:57
Hi,
This is version two of my patch series.
The use case is to be able to configure an HTTP proxy for all
subdomains of a domain where there are hundreds of subdomains.
Previously, I have been using complete regular expressions with
an escape-mechanism to match the configuration key's URLs.
According to Junio's comments, I changed this mechanism to a much
simpler one, where the user is only allowed to use globbing for
the host part of the URL. That is a user can now specify a key
`http.https://*.example.com` to match all sub-domains of
`example.com`. For now I've decided to implement it such that a
single `*` matches a single subdomain only, so for example
`https://foo.bar.example.com` would not match in this case. This
is similar to how shell-globbing works usually, so it should not
be of much surprise. It's also highlighted in the documentation.
I did not include an interdiff as too much has changed between
the two versions.
Regards
Patrick
Patrick Steinhardt (4):
mailmap: add Patrick Steinhardt's work address
urlmatch: enable normalization of URLs with globs
urlmatch: split host and port fields in `struct url_info`
urlmatch: allow globbing for the URL host part
.mailmap | 1 +
Documentation/config.txt | 5 +++-
t/t1300-repo-config.sh | 36 +++++++++++++++++++++++++
urlmatch.c | 68 +++++++++++++++++++++++++++++++++++++++++-------
urlmatch.h | 9 ++++---
5 files changed, 104 insertions(+), 15 deletions(-)
--
2.11.0
From: Patrick Steinhardt <hidden> Date: 2017-01-24 17:01:06
The `url_normalize` function is used to validate and normalize URLs. As
such, it does not allow for some special characters to be part of the
URLs that are to be normalized. As we want to allow using globs in some
configuration keys making use of URLs, namely `http.<url>.<key>`, but
still normalize them, we need to somehow enable some additional allowed
characters.
To do this without having to change all callers of `url_normalize`,
where most do not actually want globbing at all, we split off another
function `url_normalize_1`. This function accepts an additional
parameter `allow_globs`, which is subsequently called by `url_normalize`
with `allow_globs=0`.
As of now, this function is not used with globbing enabled. A caller
will be added in the following commit.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
From: Patrick Steinhardt <hidden> Date: 2017-01-24 17:01:11
The `url_info` structure contains information about a normalized URL
with the URL's components being represented by different fields. The
host and port part though are to be accessed by the same `host` field,
so that getting the host and/or port separately becomes more involved
than really necessary.
To make the port more readily accessible, split up the host and port
fields. Namely, the `host_len` will not include the port length anymore
and a new `port_off` field has been added which includes the offset to
the port, if available.
The only user of these fields is `url_normalize_1`. This change makes it
easier later on to treat host and port differently when introducing
globs for domains.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 16 ++++++++++++----
urlmatch.h | 9 +++++----
2 files changed, 17 insertions(+), 8 deletions(-)
@@ -464,11 +466,17 @@ static int match_urls(const struct url_info *url,usermatched=1;}-/* check the host and port */+/* check the host */if(url_prefix->host_len!=url->host_len||strncmp(url->url+url->host_off,url_prefix->url+url_prefix->host_off,url->host_len))-return0;/* host names and/or ports do not match */+return0;/* host names do not match */++/* check the port */+if(url_prefix->port_len!=url->port_len||+strncmp(url->url+url->port_off,+url_prefix->url+url_prefix->port_off,url->port_len))+return0;/* ports do not match *//* check the path */pathmatchlen=url_match_prefix(
@@ -18,11 +18,12 @@ struct url_info {size_tpasswd_len;/* length of passwd; if passwd_off != 0 butpasswd_len==0,anemptypasswdwasgiven*/size_thost_off;/* offset into url to start of host name (0 => none) */-size_thost_len;/* length of host name; this INCLUDES any ':portnum';+size_thost_len;/* length of host name;*fileurlsmayhavehost_len==0*/-size_tport_len;/* if a portnum is present (port_len != 0), it has-*thislength(excludingtheleading':')atthe-*endofthehostname(always0forfileurls)*/+size_tport_off;/* offset into url to start of port number (0 => none) */+size_tport_len;/* if a portnum is present (port_off != 0), it has+*thislength(excludingtheleading':')starting+*fromport_off(always0forfileurls)*/size_tpath_off;/* offset into url to the start of the url path;*thiswillalwayspointtoa'/'character*aftertheurlhasbeennormalized*/
From: Patrick Steinhardt <hidden> Date: 2017-01-24 17:01:13
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
Signed-off-by: Patrick Steinhardt <redacted>
---
Documentation/config.txt | 5 ++++-
t/t1300-repo-config.sh | 36 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 38 ++++++++++++++++++++++++++++++++++----
3 files changed, 74 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL. . Host/domain name (e.g., `example.com` in `https://example.com/`).- This field must match exactly between the config key and the URL.+ This field must match between the config key and the URL. It is+ possible to use globs in the config key to match all subdomains, e.g.+ `https://*.example.com/` to match all subdomains of `example.com`. Note+ that a glob only every matches a single part of the hostname. . Port number (e.g., `8080` in `http://example.com:8080/`). This field must match exactly between the config key and the URL.
@@ -1177,6 +1177,42 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'+test_expect_success'glob-based urlmatch''+cat>.git/config<<-\EOF&&+[http]+sslVerify+[http"https://*.example.com"]+sslVerify=false+cookieFile=/tmp/cookie.txt+EOF++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good.example.com>actual&&+test_must_be_emptyactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://good-example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://deep.nested.example.com>actual&&+test_cmpexpectactual&&++echofalse>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://good.example.com>actual&&+test_cmpexpectactual&&++{+echohttp.cookiefile/tmp/cookie.txt&&+echohttp.sslverifyfalse+}>expect&&+gitconfig--get-urlmatchHTTPhttps://good.example.com>actual&&+test_cmpexpectactual+'+# good section hygiene test_expect_failure'unsetting the last key in a section removes header''cat>.git/config<<-\EOF&&
@@ -63,6 +63,38 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticintmatch_host(conststructurl_info*url_info,+conststructurl_info*pattern_info)+{+char*url=xmemdupz(url_info->url+url_info->host_off,url_info->host_len);+char*pat=xmemdupz(pattern_info->url+pattern_info->host_off,pattern_info->host_len);+char*url_tok,*pat_tok,*url_save,*pat_save;+intmatching;++url_tok=strtok_r(url,".",&url_save);+pat_tok=strtok_r(pat,".",&pat_save);++for(;url_tok&&pat_tok;url_tok=strtok_r(NULL,".",&url_save),+pat_tok=strtok_r(NULL,".",&pat_save)){+if(!strcmp(pat_tok,"*"))+continue;/* a simple glob matches everything */++if(strcmp(url_tok,pat_tok)){+/* subdomains do not match */+matching=0;+break;+}+}++/* matching if both URL and pattern are at their ends */+matching=(url_tok==NULL&&pat_tok==NULL);++free(url);+free(pat);++returnmatching;+}+staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs){/*
@@ -467,9 +499,7 @@ static int match_urls(const struct url_info *url,}/* check the host */-if(url_prefix->host_len!=url->host_len||-strncmp(url->url+url->host_off,-url_prefix->url+url_prefix->host_off,url->host_len))+if(!match_host(url,url_prefix))return0;/* host names do not match *//* check the port */
From: Philip Oakley <hidden> Date: 2017-01-24 17:52:45
From: "Patrick Steinhardt" <redacted>
a quick comment on the documentation part ..
quoted hunk
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
Signed-off-by: Patrick Steinhardt <redacted>
---
Documentation/config.txt | 5 ++++-
t/t1300-repo-config.sh | 36 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 38 ++++++++++++++++++++++++++++++++++----
3 files changed, 74 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL.
. Host/domain name (e.g., `example.com` in `https://example.com/`).
- This field must match exactly between the config key and the URL.
+ This field must match between the config key and the URL. It is
+ possible to use globs in the config key to match all subdomains, e.g.
+ `https://*.example.com/` to match all subdomains of `example.com`. Note
+ that a glob only every matches a single part of the hostname.
[s/every/ever/ ?]
the "match all subdomains" appears to contradict the "a glob only ever
matches a single part ".
Maybe borrow the example from the 0/4 cover letter
"so for example `https://foo.bar.example.com` would not match in the case of
`http.https://*.example.com` " (If I understood it correctly.
A simple example often clarifies much better than more words.
--
Philip
quoted hunk
. Port number (e.g., `8080` in `http://example.com:8080/`).
This field must match exactly between the config key and the URL.
@@ -467,9 +499,7 @@ static int match_urls(const struct url_info *url, } /* check the host */- if (url_prefix->host_len != url->host_len ||- strncmp(url->url + url->host_off,- url_prefix->url + url_prefix->host_off, url->host_len))+ if (!match_host(url, url_prefix)) return 0; /* host names do not match */ /* check the port */
@@ -512,7 +542,7 @@ int urlmatch_config_entry(const char *var, const char
From: Patrick Steinhardt <hidden> Date: 2017-01-25 09:56:59
The `url_normalize` function is used to validate and normalize URLs. As
such, it does not allow for some special characters to be part of the
URLs that are to be normalized. As we want to allow using globs in some
configuration keys making use of URLs, namely `http.<url>.<key>`, but
still normalize them, we need to somehow enable some additional allowed
characters.
To do this without having to change all callers of `url_normalize`,
where most do not actually want globbing at all, we split off another
function `url_normalize_1`. This function accepts an additional
parameter `allow_globs`, which is subsequently called by `url_normalize`
with `allow_globs=0`.
As of now, this function is not used with globbing enabled. A caller
will be added in the following commit.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
From: Patrick Steinhardt <hidden> Date: 2017-01-25 09:57:02
Hi,
This is version three of my patch series. The previous version
can be found at [1]. The use case is to be able to configure an
HTTP proxy for all subdomains of a domain where there are
hundreds of subdomains.
This version addresses a comment by Philip Oakley regarding the
documentation. You can find the interdiff below.
Regards
Patrick
[1]: http://public-inbox.org/git/20170124170031.18069-1-patrick.steinhardt@elego.de/T/#u
Patrick Steinhardt (4):
mailmap: add Patrick Steinhardt's work address
urlmatch: enable normalization of URLs with globs
urlmatch: split host and port fields in `struct url_info`
urlmatch: allow globbing for the URL host part
.mailmap | 1 +
Documentation/config.txt | 5 +++-
t/t1300-repo-config.sh | 36 +++++++++++++++++++++++++
urlmatch.c | 68 +++++++++++++++++++++++++++++++++++++++++-------
urlmatch.h | 9 ++++---
5 files changed, 104 insertions(+), 15 deletions(-)
@@ -1915,9 +1915,9 @@ http.<url>.*:: . Host/domain name (e.g., `example.com` in `https://example.com/`). This field must match between the config key and the URL. It is- possible to use globs in the config key to match all subdomains, e.g.- `https://*.example.com/` to match all subdomains of `example.com`. Note- that a glob only every matches a single part of the hostname.+ possible to specify a `*` as part of the host name to match all subdomains+ at this level. `https://*.example.com/` for example would match+ `https://foo.example.com/`, but not `https://foo.bar.example.com/`. . Port number (e.g., `8080` in `http://example.com:8080/`). This field must match exactly between the config key and the URL.
From: Patrick Steinhardt <hidden> Date: 2017-01-25 09:57:05
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
Signed-off-by: Patrick Steinhardt <redacted>
---
Documentation/config.txt | 5 ++++-
t/t1300-repo-config.sh | 36 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 38 ++++++++++++++++++++++++++++++++++----
3 files changed, 74 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL. . Host/domain name (e.g., `example.com` in `https://example.com/`).- This field must match exactly between the config key and the URL.+ This field must match between the config key and the URL. It is+ possible to specify a `*` as part of the host name to match all subdomains+ at this level. `https://*.example.com/` for example would match+ `https://foo.example.com/`, but not `https://foo.bar.example.com/`. . Port number (e.g., `8080` in `http://example.com:8080/`). This field must match exactly between the config key and the URL.
@@ -1177,6 +1177,42 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'+test_expect_success'glob-based urlmatch''+cat>.git/config<<-\EOF&&+[http]+sslVerify+[http"https://*.example.com"]+sslVerify=false+cookieFile=/tmp/cookie.txt+EOF++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good.example.com>actual&&+test_must_be_emptyactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://good-example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://deep.nested.example.com>actual&&+test_cmpexpectactual&&++echofalse>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://good.example.com>actual&&+test_cmpexpectactual&&++{+echohttp.cookiefile/tmp/cookie.txt&&+echohttp.sslverifyfalse+}>expect&&+gitconfig--get-urlmatchHTTPhttps://good.example.com>actual&&+test_cmpexpectactual+'+# good section hygiene test_expect_failure'unsetting the last key in a section removes header''cat>.git/config<<-\EOF&&
@@ -63,6 +63,38 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticintmatch_host(conststructurl_info*url_info,+conststructurl_info*pattern_info)+{+char*url=xmemdupz(url_info->url+url_info->host_off,url_info->host_len);+char*pat=xmemdupz(pattern_info->url+pattern_info->host_off,pattern_info->host_len);+char*url_tok,*pat_tok,*url_save,*pat_save;+intmatching;++url_tok=strtok_r(url,".",&url_save);+pat_tok=strtok_r(pat,".",&pat_save);++for(;url_tok&&pat_tok;url_tok=strtok_r(NULL,".",&url_save),+pat_tok=strtok_r(NULL,".",&pat_save)){+if(!strcmp(pat_tok,"*"))+continue;/* a simple glob matches everything */++if(strcmp(url_tok,pat_tok)){+/* subdomains do not match */+matching=0;+break;+}+}++/* matching if both URL and pattern are at their ends */+matching=(url_tok==NULL&&pat_tok==NULL);++free(url);+free(pat);++returnmatching;+}+staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs){/*
@@ -467,9 +499,7 @@ static int match_urls(const struct url_info *url,}/* check the host */-if(url_prefix->host_len!=url->host_len||-strncmp(url->url+url->host_off,-url_prefix->url+url_prefix->host_off,url->host_len))+if(!match_host(url,url_prefix))return0;/* host names do not match *//* check the port */
From: Patrick Steinhardt <hidden> Date: 2017-01-25 09:57:12
The `url_info` structure contains information about a normalized URL
with the URL's components being represented by different fields. The
host and port part though are to be accessed by the same `host` field,
so that getting the host and/or port separately becomes more involved
than really necessary.
To make the port more readily accessible, split up the host and port
fields. Namely, the `host_len` will not include the port length anymore
and a new `port_off` field has been added which includes the offset to
the port, if available.
The only user of these fields is `url_normalize_1`. This change makes it
easier later on to treat host and port differently when introducing
globs for domains.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 16 ++++++++++++----
urlmatch.h | 9 +++++----
2 files changed, 17 insertions(+), 8 deletions(-)
@@ -464,11 +466,17 @@ static int match_urls(const struct url_info *url,usermatched=1;}-/* check the host and port */+/* check the host */if(url_prefix->host_len!=url->host_len||strncmp(url->url+url->host_off,url_prefix->url+url_prefix->host_off,url->host_len))-return0;/* host names and/or ports do not match */+return0;/* host names do not match */++/* check the port */+if(url_prefix->port_len!=url->port_len||+strncmp(url->url+url->port_off,+url_prefix->url+url_prefix->port_off,url->port_len))+return0;/* ports do not match *//* check the path */pathmatchlen=url_match_prefix(
@@ -18,11 +18,12 @@ struct url_info {size_tpasswd_len;/* length of passwd; if passwd_off != 0 butpasswd_len==0,anemptypasswdwasgiven*/size_thost_off;/* offset into url to start of host name (0 => none) */-size_thost_len;/* length of host name; this INCLUDES any ':portnum';+size_thost_len;/* length of host name;*fileurlsmayhavehost_len==0*/-size_tport_len;/* if a portnum is present (port_len != 0), it has-*thislength(excludingtheleading':')atthe-*endofthehostname(always0forfileurls)*/+size_tport_off;/* offset into url to start of port number (0 => none) */+size_tport_len;/* if a portnum is present (port_off != 0), it has+*thislength(excludingtheleading':')starting+*fromport_off(always0forfileurls)*/size_tpath_off;/* offset into url to the start of the url path;*thiswillalwayspointtoa'/'character*aftertheurlhasbeennormalized*/
From: Patrick Steinhardt <hidden> Date: 2017-01-25 09:58:07
On Tue, Jan 24, 2017 at 05:52:39PM -0000, Philip Oakley wrote:
From: "Patrick Steinhardt" <redacted>
a quick comment on the documentation part ..
quoted
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
Signed-off-by: Patrick Steinhardt <redacted>
---
Documentation/config.txt | 5 ++++-
t/t1300-repo-config.sh | 36 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 38 ++++++++++++++++++++++++++++++++++----
3 files changed, 74 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL.
. Host/domain name (e.g., `example.com` in `https://example.com/`).
- This field must match exactly between the config key and the URL.
+ This field must match between the config key and the URL. It is
+ possible to use globs in the config key to match all subdomains, e.g.
+ `https://*.example.com/` to match all subdomains of `example.com`. Note
+ that a glob only every matches a single part of the hostname.
[s/every/ever/ ?]
the "match all subdomains" appears to contradict the "a glob only ever
matches a single part ".
Maybe borrow the example from the 0/4 cover letter
"so for example `https://foo.bar.example.com` would not match in the case of
`http.https://*.example.com` " (If I understood it correctly.
A simple example often clarifies much better than more words.
Thanks! (Hopefully) improved this in v3.
Regards
Patrick
From: Junio C Hamano <hidden> Date: 2017-01-26 20:47:44
Patrick Steinhardt [off-list ref] writes:
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
This is probably a useful improvement.
Having said that, when I mentioned "glob", I meant to also support
something like this:
https://www[1-4].ibm.com/
And when people read "glob", that is what they expect.
So calling this "the ability to use globbing" is misleading.
The last paragraph in the log message above needs a bit of
tweaking, perhaps like this:
Allow users to write an asterisk '*' in place of any 'host'
or 'subdomain' label as part of the host name. For example,
"http.https://*.example.com.proxy" sets "http.proxy" for all
direct subdomains of "https://example.com",
e.g. "https://foo.example.com", but not
"https://foo.bar.example.com".
Fortunately, your update to config.txt, which is facing the end
users, does not misuse the word and instead is explicit that the
only thing the matcher does is to match '*' to a single hierarchy.
It is clear that even http://www*.ibm.com/ is not supported from
the description, which is good.
. Host/domain name (e.g., `example.com` in `https://example.com/`).
- This field must match exactly between the config key and the URL.
+ This field must match between the config key and the URL. It is
+ possible to specify a `*` as part of the host name to match all subdomains
+ at this level. `https://*.example.com/` for example would match
+ `https://foo.example.com/`, but not `https://foo.bar.example.com/`.
This is good as-is.
quoted hunk
. Port number (e.g., `8080` in `http://example.com:8080/`).
This field must match exactly between the config key and the URL.
Hmph, this will be the first use of strtok_r() in our codebase.
Does everybody have it?
For a use like this where your delimiter set is a singleton, it may
be simpler to do the usual strchrnul() or memchr() based loop. The
attached is my attempt to do so on top of this patch.
@@ -63,36 +63,47 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticconstchar*end_of_token(constchar*s,intc,size_tn)+{+constchar*next=memchr(s,c,n);+if(!next)+next=s+n;+returnnext;+}+staticintmatch_host(conststructurl_info*url_info,conststructurl_info*pattern_info){-char*url=xmemdupz(url_info->url+url_info->host_off,url_info->host_len);-char*pat=xmemdupz(pattern_info->url+pattern_info->host_off,pattern_info->host_len);-char*url_tok,*pat_tok,*url_save,*pat_save;-intmatching;--url_tok=strtok_r(url,".",&url_save);-pat_tok=strtok_r(pat,".",&pat_save);--for(;url_tok&&pat_tok;url_tok=strtok_r(NULL,".",&url_save),-pat_tok=strtok_r(NULL,".",&pat_save)){-if(!strcmp(pat_tok,"*"))-continue;/* a simple glob matches everything */--if(strcmp(url_tok,pat_tok)){-/* subdomains do not match */-matching=0;-break;-}+constchar*url=url_info->url+url_info->host_off;+constchar*pat=pattern_info->url+pattern_info->host_off;+inturl_len=url_info->host_len;+intpat_len=pattern_info->host_len;++while(url_len&&pat_len){+constchar*url_next=end_of_token(url,'.',url_len);+constchar*pat_next=end_of_token(pat,'.',pat_len);++if(pat_next==pat+1&&pat[0]=='*')+/* wildcard matches anything */+;+elseif((pat_next-pat)==(url_next-url)&&+!memcmp(url,pat,url_next-url))+/* the components are the same */+;+else+return0;/* found an unmatch */++if(url_next<url+url_len)+url_next++;+url_len-=url_next-url;+url=url_next;+if(pat_next<pat+pat_len)+pat_next++;+pat_len-=pat_next-pat;+pat=pat_next;}-/* matching if both URL and pattern are at their ends */-matching=(url_tok==NULL&&pat_tok==NULL);--free(url);-free(pat);--returnmatching;+return1;}staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs)
From: Patrick Steinhardt <hidden> Date: 2017-01-27 06:21:56
On Thu, Jan 26, 2017 at 12:43:31PM -0800, Junio C Hamano wrote:
Patrick Steinhardt [off-list ref] writes:
quoted
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
This commit introduces the ability to use globbing in the host-part of
the URLs. A user can simply specify a `*` as part of the host name to
match all subdomains at this level. For example adding a configuration
key `http.https://*.example.com.proxy` will match all subdomains of
`https://example.com`.
This is probably a useful improvement.
Having said that, when I mentioned "glob", I meant to also support
something like this:
https://www[1-4].ibm.com/
The problem with additional extended syntax like proposed by you
is that we would indeed need an escaping mechanism here. '[]' are
already allowed inside the host part to enable IPv6 hosts of the
form 'https://[2001:0db8:]/', so the syntax is now ambiguous. So
we have to be cautios which characters to enable for globbing
syntax. As of now, I think we can only safely include '*' and '?'
here without escaping mechanisms.
If additional use cases come up we might still extend the syntax
later on to allow for more special syntax.
And when people read "glob", that is what they expect.
So calling this "the ability to use globbing" is misleading.
The last paragraph in the log message above needs a bit of
tweaking, perhaps like this:
Allow users to write an asterisk '*' in place of any 'host'
or 'subdomain' label as part of the host name. For example,
"http.https://*.example.com.proxy" sets "http.proxy" for all
direct subdomains of "https://example.com",
e.g. "https://foo.example.com", but not
"https://foo.bar.example.com".
Fortunately, your update to config.txt, which is facing the end
users, does not misuse the word and instead is explicit that the
only thing the matcher does is to match '*' to a single hierarchy.
It is clear that even http://www*.ibm.com/ is not supported from
the description, which is good.
I agree that globbing is the wrong word here. I'll swap in
"wildcard" where applicable.
I'll send a version 4 later on. Thanks again for your feedback
and improvements.
Regards
Patrick
From: Patrick Steinhardt <hidden> Date: 2017-01-27 10:35:35
The `url_normalize` function is used to validate and normalize URLs. As
such, it does not allow for some special characters to be part of the
URLs that are to be normalized. As we want to allow using globs in some
configuration keys making use of URLs, namely `http.<url>.<key>`, but
still normalize them, we need to somehow enable some additional allowed
characters.
To do this without having to change all callers of `url_normalize`,
where most do not actually want globbing at all, we split off another
function `url_normalize_1`. This function accepts an additional
parameter `allow_globs`, which is subsequently called by `url_normalize`
with `allow_globs=0`.
As of now, this function is not used with globbing enabled. A caller
will be added in the following commit.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
From: Patrick Steinhardt <hidden> Date: 2017-01-27 10:35:38
Hi,
so this is part four of my patch series. The previous version can
be found at [1]. The use case is to be able to configure an HTTP
proxy for all subdomains of a domain where there are hundreds of
subdomains.
Changes to the previous version:
- applied Junio's proposed patch to replace `strtok_r` with a
`memchr`-based loop
- applied Junio's proposed rewrite of the commit message of
patch 5
- I realized that with my patches, "ranking" of URLs was broken.
Previously, we've always taken the longest matching URL. As
previously, only the user and path could actually differ, only
these two components were used for the comparison. I've
changed this now to also include the host part so that URLs
with a longer host will take precedence. This resulted in a
the patch 4.
- New tests are included which examine if the precedence-rules
are actually honored correctly. The tests are part of patches
4 and 5.
You can find the interdiff below.
Regards
Patrick
[1]: http://public-inbox.org/git/20170125095648.4116-1-patrick.steinhardt@elego.de/T/#t
Patrick Steinhardt (5):
mailmap: add Patrick Steinhardt's work address
urlmatch: enable normalization of URLs with globs
urlmatch: split host and port fields in `struct url_info`
urlmatch: include host and port in urlmatch length
urlmatch: allow globbing for the URL host part
.mailmap | 1 +
Documentation/config.txt | 5 +-
t/t1300-repo-config.sh | 105 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 138 +++++++++++++++++++++++++++++++++++------------
urlmatch.h | 12 +++--
5 files changed, 220 insertions(+), 41 deletions(-)
--
2.11.0
@@ -1177,7 +1177,72 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'-test_expect_success'glob-based urlmatch''+test_expect_success'urlmatch favors more specific URLs''+cat>.git/config<<-\EOF&&+[http"https://example.com/"]+cookieFile=/tmp/root.txt+[http"https://example.com/subdirectory"]+cookieFile=/tmp/subdirectory.txt+[http"https://user@example.com/"]+cookieFile=/tmp/user.txt+[http"https://averylonguser@example.com/"]+cookieFile=/tmp/averylonguser.txt+[http"https://preceding.example.com"]+cookieFile=/tmp/preceding.txt+[http"https://*.example.com"]+cookieFile=/tmp/wildcard.txt+[http"https://*.example.com/wildcardwithsubdomain"]+cookieFile=/tmp/wildcardwithsubdomain.txt+[http"https://trailing.example.com"]+cookieFile=/tmp/trailing.txt+[http"https://user@*.example.com/"]+cookieFile=/tmp/wildcardwithuser.txt+[http"https://sub.example.com/"]+cookieFile=/tmp/sub.txt+EOF++echohttp.cookiefile/tmp/root.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory/nested>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/user.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@example.com/>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://averylonguser@example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/preceding.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://preceding.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/wildcard.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://wildcard.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://sub.example.com/wildcardwithsubdomain>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/trailing.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://trailing.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@sub.example.com>actual&&+test_cmpexpectactual+'++test_expect_success'urlmatch with wildcard''cat>.git/config<<-\EOF&&[http]sslVerify
@@ -63,36 +63,47 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticconstchar*end_of_token(constchar*s,intc,size_tn)+{+constchar*next=memchr(s,c,n);+if(!next)+next=s+n;+returnnext;+}+staticintmatch_host(conststructurl_info*url_info,conststructurl_info*pattern_info){-char*url=xmemdupz(url_info->url+url_info->host_off,url_info->host_len);-char*pat=xmemdupz(pattern_info->url+pattern_info->host_off,pattern_info->host_len);-char*url_tok,*pat_tok,*url_save,*pat_save;-intmatching;+constchar*url=url_info->url+url_info->host_off;+constchar*pat=pattern_info->url+pattern_info->host_off;+inturl_len=url_info->host_len;+intpat_len=pattern_info->host_len;-url_tok=strtok_r(url,".",&url_save);-pat_tok=strtok_r(pat,".",&pat_save);+while(url_len&&pat_len){+constchar*url_next=end_of_token(url,'.',url_len);+constchar*pat_next=end_of_token(pat,'.',pat_len);-for(;url_tok&&pat_tok;url_tok=strtok_r(NULL,".",&url_save),-pat_tok=strtok_r(NULL,".",&pat_save)){-if(!strcmp(pat_tok,"*"))-continue;/* a simple glob matches everything */+if(pat_next==pat+1&&pat[0]=='*')+/* wildcard matches anything */+;+elseif((pat_next-pat)==(url_next-url)&&+!memcmp(url,pat,url_next-url))+/* the components are the same */+;+else+return0;/* found an unmatch */-if(strcmp(url_tok,pat_tok)){-/* subdomains do not match */-matching=0;-break;-}+if(url_next<url+url_len)+url_next++;+url_len-=url_next-url;+url=url_next;+if(pat_next<pat+pat_len)+pat_next++;+pat_len-=pat_next-pat;+pat=pat_next;}-/* matching if both URL and pattern are at their ends */-matching=(url_tok==NULL&&pat_tok==NULL);--free(url);-free(pat);--returnmatching;+return(!url_len&&!pat_len);}staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs)
@@ -477,8 +488,8 @@ static int match_urls(const struct url_info *url,*containedausernameorfalseifurl_prefixdidnothavea*username.Ifthereisnomatch*exactusermatchisleftuntouched.*/-intusermatched=0;-intpathmatchlen;+charusermatched=0;+size_tpathmatchlen;if(!url||!url_prefix||!url->url||!url_prefix->url)return0;
@@ -513,22 +524,38 @@ static int match_urls(const struct url_info *url,url->url+url->path_off,url_prefix->url+url_prefix->path_off,url_prefix->url_len-url_prefix->path_off);+if(!pathmatchlen)+return0;/* paths do not match */-if(pathmatchlen&&exactusermatch)-*exactusermatch=usermatched;-returnpathmatchlen;+if(match){+match->hostmatch_len=url_prefix->host_len;+match->pathmatch_len=pathmatchlen;+match->user_matched=usermatched;+}++return1;+}++staticintcmp_matches(conststructurlmatch_item*a,+conststructurlmatch_item*b)+{+if(a->hostmatch_len!=b->hostmatch_len)+returna->hostmatch_len<b->hostmatch_len?-1:1;+if(a->pathmatch_len!=b->pathmatch_len)+returna->pathmatch_len<b->pathmatch_len?-1:1;+if(a->user_matched!=b->user_matched)+returnb->user_matched?-1:1;+return0;}inturlmatch_config_entry(constchar*var,constchar*value,void*cb){structstring_list_item*item;structurlmatch_config*collect=cb;-structurlmatch_item*matched;+structurlmatch_itemmatched;structurl_info*url=&collect->url;constchar*key,*dot;structstrbufsynthkey=STRBUF_INIT;-size_tmatched_len=0;-intuser_matched=0;intretval;if(!skip_prefix(var,collect->section,&key)||*(key++)!='.'){
@@ -558,24 +585,17 @@ int urlmatch_config_entry(const char *var, const char *value, void *cb)item=string_list_insert(&collect->vars,key);if(!item->util){-matched=xcalloc(1,sizeof(*matched));-item->util=matched;+item->util=xcalloc(1,sizeof(matched));}else{-matched=item->util;-/*-*Isourmatchshorter?Isourmatchthesame-*length,andwithoutuserwhilethecurrent-*candidateiswithuser?Thenwecannotuseit.-*/-if(matched_len<matched->matched_len||-((matched_len==matched->matched_len)&&-(!user_matched&&matched->user_matched)))+if(cmp_matches(&matched,item->util)<=0)+/*+*Ourmatchisworsethantheoldone,+*wecannotuseit.+*/return0;-/* Otherwise, replace it with this one. */}-matched->matched_len=matched_len;-matched->user_matched=user_matched;+memcpy(item->util,&matched,sizeof(matched));strbuf_addstr(&synthkey,collect->section);strbuf_addch(&synthkey,'.');strbuf_addstr(&synthkey,key);
From: Patrick Steinhardt <hidden> Date: 2017-01-27 10:35:40
In order to be able to rank positive matches by `urlmatch`, we inspect
the path length and user part to decide whether a match is better than
another match. As all other parts are matched exactly between both URLs,
this is the right thing to do right now.
In the future, though, we want to introduce wild cards for the domain
part. When doing this, though, it does not make sense anymore to only
compare the path lengths. Instead, we also want to compare the domain
lengths to determine which of both URLs matches the host part more
closely.
Signed-off-by: Patrick Steinhardt <redacted>
---
t/t1300-repo-config.sh | 33 ++++++++++++++++++++++++++++
urlmatch.c | 59 +++++++++++++++++++++++++++++---------------------
urlmatch.h | 3 ++-
3 files changed, 69 insertions(+), 26 deletions(-)
@@ -1177,6 +1177,39 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'+test_expect_success'urlmatch favors more specific URLs''+cat>.git/config<<-\EOF&&+[http"https://example.com/"]+cookieFile=/tmp/root.txt+[http"https://example.com/subdirectory"]+cookieFile=/tmp/subdirectory.txt+[http"https://user@example.com/"]+cookieFile=/tmp/user.txt+[http"https://averylonguser@example.com/"]+cookieFile=/tmp/averylonguser.txt+EOF++echohttp.cookiefile/tmp/root.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory/nested>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/user.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@example.com/>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://averylonguser@example.com/subdirectory>actual&&+test_cmpexpectactual+'+# good section hygiene test_expect_failure'unsetting the last key in a section removes header''cat>.git/config<<-\EOF&&
@@ -445,8 +445,8 @@ static int match_urls(const struct url_info *url,*containedausernameorfalseifurl_prefixdidnothavea*username.Ifthereisnomatch*exactusermatchisleftuntouched.*/-intusermatched=0;-intpathmatchlen;+charusermatched=0;+size_tpathmatchlen;if(!url||!url_prefix||!url->url||!url_prefix->url)return0;
@@ -483,22 +483,38 @@ static int match_urls(const struct url_info *url,url->url+url->path_off,url_prefix->url+url_prefix->path_off,url_prefix->url_len-url_prefix->path_off);+if(!pathmatchlen)+return0;/* paths do not match */-if(pathmatchlen&&exactusermatch)-*exactusermatch=usermatched;-returnpathmatchlen;+if(match){+match->hostmatch_len=url_prefix->host_len;+match->pathmatch_len=pathmatchlen;+match->user_matched=usermatched;+}++return1;+}++staticintcmp_matches(conststructurlmatch_item*a,+conststructurlmatch_item*b)+{+if(a->hostmatch_len!=b->hostmatch_len)+returna->hostmatch_len<b->hostmatch_len?-1:1;+if(a->pathmatch_len!=b->pathmatch_len)+returna->pathmatch_len<b->pathmatch_len?-1:1;+if(a->user_matched!=b->user_matched)+returnb->user_matched?-1:1;+return0;}inturlmatch_config_entry(constchar*var,constchar*value,void*cb){structstring_list_item*item;structurlmatch_config*collect=cb;-structurlmatch_item*matched;+structurlmatch_itemmatched;structurl_info*url=&collect->url;constchar*key,*dot;structstrbufsynthkey=STRBUF_INIT;-size_tmatched_len=0;-intuser_matched=0;intretval;if(!skip_prefix(var,collect->section,&key)||*(key++)!='.'){
@@ -528,24 +544,17 @@ int urlmatch_config_entry(const char *var, const char *value, void *cb)item=string_list_insert(&collect->vars,key);if(!item->util){-matched=xcalloc(1,sizeof(*matched));-item->util=matched;+item->util=xcalloc(1,sizeof(matched));}else{-matched=item->util;-/*-*Isourmatchshorter?Isourmatchthesame-*length,andwithoutuserwhilethecurrent-*candidateiswithuser?Thenwecannotuseit.-*/-if(matched_len<matched->matched_len||-((matched_len==matched->matched_len)&&-(!user_matched&&matched->user_matched)))+if(cmp_matches(&matched,item->util)<=0)+/*+*Ourmatchisworsethantheoldone,+*wecannotuseit.+*/return0;-/* Otherwise, replace it with this one. */}-matched->matched_len=matched_len;-matched->user_matched=user_matched;+memcpy(item->util,&matched,sizeof(matched));strbuf_addstr(&synthkey,collect->section);strbuf_addch(&synthkey,'.');strbuf_addstr(&synthkey,key);
From: Patrick Steinhardt <hidden> Date: 2017-01-27 10:35:43
The `url_info` structure contains information about a normalized URL
with the URL's components being represented by different fields. The
host and port part though are to be accessed by the same `host` field,
so that getting the host and/or port separately becomes more involved
than really necessary.
To make the port more readily accessible, split up the host and port
fields. Namely, the `host_len` will not include the port length anymore
and a new `port_off` field has been added which includes the offset to
the port, if available.
The only user of these fields is `url_normalize_1`. This change makes it
easier later on to treat host and port differently when introducing
globs for domains.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 16 ++++++++++++----
urlmatch.h | 9 +++++----
2 files changed, 17 insertions(+), 8 deletions(-)
@@ -464,11 +466,17 @@ static int match_urls(const struct url_info *url,usermatched=1;}-/* check the host and port */+/* check the host */if(url_prefix->host_len!=url->host_len||strncmp(url->url+url->host_off,url_prefix->url+url_prefix->host_off,url->host_len))-return0;/* host names and/or ports do not match */+return0;/* host names do not match */++/* check the port */+if(url_prefix->port_len!=url->port_len||+strncmp(url->url+url->port_off,+url_prefix->url+url_prefix->port_off,url->port_len))+return0;/* ports do not match *//* check the path */pathmatchlen=url_match_prefix(
@@ -18,11 +18,12 @@ struct url_info {size_tpasswd_len;/* length of passwd; if passwd_off != 0 butpasswd_len==0,anemptypasswdwasgiven*/size_thost_off;/* offset into url to start of host name (0 => none) */-size_thost_len;/* length of host name; this INCLUDES any ':portnum';+size_thost_len;/* length of host name;*fileurlsmayhavehost_len==0*/-size_tport_len;/* if a portnum is present (port_len != 0), it has-*thislength(excludingtheleading':')atthe-*endofthehostname(always0forfileurls)*/+size_tport_off;/* offset into url to start of port number (0 => none) */+size_tport_len;/* if a portnum is present (port_off != 0), it has+*thislength(excludingtheleading':')starting+*fromport_off(always0forfileurls)*/size_tpath_off;/* offset into url to the start of the url path;*thiswillalwayspointtoa'/'character*aftertheurlhasbeennormalized*/
From: Patrick Steinhardt <hidden> Date: 2017-01-27 10:35:45
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
Allow users to write an asterisk '*' in place of any 'host' or
'subdomain' label as part of the host name. For example,
"http.https://*.example.com.proxy" sets "http.proxy" for all direct
subdomains of "https://example.com", e.g. "https://foo.example.com", but
not "https://foo.bar.example.com".
Signed-off-by: Patrick Steinhardt <redacted>
Helped-by: Junio C Hamano [off-list ref]
---
Documentation/config.txt | 5 +++-
t/t1300-repo-config.sh | 72 ++++++++++++++++++++++++++++++++++++++++++++++++
urlmatch.c | 49 +++++++++++++++++++++++++++++---
3 files changed, 121 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL. . Host/domain name (e.g., `example.com` in `https://example.com/`).- This field must match exactly between the config key and the URL.+ This field must match between the config key and the URL. It is+ possible to specify a `*` as part of the host name to match all subdomains+ at this level. `https://*.example.com/` for example would match+ `https://foo.example.com/`, but not `https://foo.bar.example.com/`. . Port number (e.g., `8080` in `http://example.com:8080/`). This field must match exactly between the config key and the URL.
@@ -1187,6 +1187,18 @@ test_expect_success 'urlmatch favors more specific URLs' 'cookieFile=/tmp/user.txt[http"https://averylonguser@example.com/"]cookieFile=/tmp/averylonguser.txt+[http"https://preceding.example.com"]+cookieFile=/tmp/preceding.txt+[http"https://*.example.com"]+cookieFile=/tmp/wildcard.txt+[http"https://*.example.com/wildcardwithsubdomain"]+cookieFile=/tmp/wildcardwithsubdomain.txt+[http"https://trailing.example.com"]+cookieFile=/tmp/trailing.txt+[http"https://user@*.example.com/"]+cookieFile=/tmp/wildcardwithuser.txt+[http"https://sub.example.com/"]+cookieFile=/tmp/sub.txtEOFechohttp.cookiefile/tmp/root.txt>expect&&
@@ -1207,6 +1219,66 @@ test_expect_success 'urlmatch favors more specific URLs' 'echohttp.cookiefile/tmp/subdirectory.txt>expect&&gitconfig--get-urlmatchHTTPhttps://averylonguser@example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/preceding.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://preceding.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/wildcard.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://wildcard.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://sub.example.com/wildcardwithsubdomain>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/trailing.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://trailing.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@sub.example.com>actual&&+test_cmpexpectactual+'++test_expect_success'urlmatch with wildcard''+cat>.git/config<<-\EOF&&+[http]+sslVerify+[http"https://*.example.com"]+sslVerify=false+cookieFile=/tmp/cookie.txt+EOF++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good.example.com>actual&&+test_must_be_emptyactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://good-example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://deep.nested.example.com>actual&&+test_cmpexpectactual&&++echofalse>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://good.example.com>actual&&+test_cmpexpectactual&&++{+echohttp.cookiefile/tmp/cookie.txt&&+echohttp.sslverifyfalse+}>expect&&+gitconfig--get-urlmatchHTTPhttps://good.example.com>actual&&+test_cmpexpectactual&&++echohttp.sslverify>expect&&+gitconfig--get-urlmatchHTTPhttps://more.example.com.au>actual&&test_cmpexpectactual'
@@ -63,6 +63,49 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticconstchar*end_of_token(constchar*s,intc,size_tn)+{+constchar*next=memchr(s,c,n);+if(!next)+next=s+n;+returnnext;+}++staticintmatch_host(conststructurl_info*url_info,+conststructurl_info*pattern_info)+{+constchar*url=url_info->url+url_info->host_off;+constchar*pat=pattern_info->url+pattern_info->host_off;+inturl_len=url_info->host_len;+intpat_len=pattern_info->host_len;++while(url_len&&pat_len){+constchar*url_next=end_of_token(url,'.',url_len);+constchar*pat_next=end_of_token(pat,'.',pat_len);++if(pat_next==pat+1&&pat[0]=='*')+/* wildcard matches anything */+;+elseif((pat_next-pat)==(url_next-url)&&+!memcmp(url,pat,url_next-url))+/* the components are the same */+;+else+return0;/* found an unmatch */++if(url_next<url+url_len)+url_next++;+url_len-=url_next-url;+url=url_next;+if(pat_next<pat+pat_len)+pat_next++;+pat_len-=pat_next-pat;+pat=pat_next;+}++return(!url_len&&!pat_len);+}+staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs){/*
@@ -467,9 +510,7 @@ static int match_urls(const struct url_info *url,}/* check the host */-if(url_prefix->host_len!=url->host_len||-strncmp(url->url+url->host_off,-url_prefix->url+url_prefix->host_off,url->host_len))+if(!match_host(url,url_prefix))return0;/* host names do not match *//* check the port */
From: Patrick Steinhardt <hidden> Date: 2017-01-31 09:05:02
From: Patrick Steinhardt <redacted>
The URL matching function computes for two URLs whether they match not.
The match is performed by splitting up the URL into different parts and
then doing an exact comparison with the to-be-matched URL.
The main user of `urlmatch` is the configuration subsystem. It allows to
set certain configurations based on the URL which is being connected to
via keys like `http.<url>.*`. A common use case for this is to set
proxies for only some remotes which match the given URL. Unfortunately,
having exact matches for all parts of the URL can become quite tedious
in some setups. Imagine for example a corporate network where there are
dozens or even hundreds of subdomains, which would have to be configured
individually.
Allow users to write an asterisk '*' in place of any 'host' or
'subdomain' label as part of the host name. For example,
"http.https://*.example.com.proxy" sets "http.proxy" for all direct
subdomains of "https://example.com", e.g. "https://foo.example.com", but
not "https://foo.bar.example.com".
Signed-off-by: Patrick Steinhardt <redacted>
Helped-by: Junio C Hamano [off-list ref]
---
Documentation/config.txt | 5 +++-
t/t1300-repo-config.sh | 72 ++++++++++++++++++++++++++++++++++++++++++++++++
urlmatch.c | 49 +++++++++++++++++++++++++++++---
3 files changed, 121 insertions(+), 5 deletions(-)
@@ -1914,7 +1914,10 @@ http.<url>.*:: must match exactly between the config key and the URL. . Host/domain name (e.g., `example.com` in `https://example.com/`).- This field must match exactly between the config key and the URL.+ This field must match between the config key and the URL. It is+ possible to specify a `*` as part of the host name to match all subdomains+ at this level. `https://*.example.com/` for example would match+ `https://foo.example.com/`, but not `https://foo.bar.example.com/`. . Port number (e.g., `8080` in `http://example.com:8080/`). This field must match exactly between the config key and the URL.
@@ -1187,6 +1187,18 @@ test_expect_success 'urlmatch favors more specific URLs' 'cookieFile=/tmp/user.txt[http"https://averylonguser@example.com/"]cookieFile=/tmp/averylonguser.txt+[http"https://preceding.example.com"]+cookieFile=/tmp/preceding.txt+[http"https://*.example.com"]+cookieFile=/tmp/wildcard.txt+[http"https://*.example.com/wildcardwithsubdomain"]+cookieFile=/tmp/wildcardwithsubdomain.txt+[http"https://trailing.example.com"]+cookieFile=/tmp/trailing.txt+[http"https://user@*.example.com/"]+cookieFile=/tmp/wildcardwithuser.txt+[http"https://sub.example.com/"]+cookieFile=/tmp/sub.txtEOFechohttp.cookiefile/tmp/root.txt>expect&&
@@ -1207,6 +1219,66 @@ test_expect_success 'urlmatch favors more specific URLs' 'echohttp.cookiefile/tmp/subdirectory.txt>expect&&gitconfig--get-urlmatchHTTPhttps://averylonguser@example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/preceding.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://preceding.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/wildcard.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://wildcard.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://sub.example.com/wildcardwithsubdomain>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/trailing.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://trailing.example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/sub.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@sub.example.com>actual&&+test_cmpexpectactual+'++test_expect_success'urlmatch with wildcard''+cat>.git/config<<-\EOF&&+[http]+sslVerify+[http"https://*.example.com"]+sslVerify=false+cookieFile=/tmp/cookie.txt+EOF++test_expect_code1gitconfig--bool--get-urlmatchdoesnt.existhttps://good.example.com>actual&&+test_must_be_emptyactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.SSLverifyhttps://good-example.com>actual&&+test_cmpexpectactual&&++echotrue>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://deep.nested.example.com>actual&&+test_cmpexpectactual&&++echofalse>expect&&+gitconfig--bool--get-urlmatchhttp.sslverifyhttps://good.example.com>actual&&+test_cmpexpectactual&&++{+echohttp.cookiefile/tmp/cookie.txt&&+echohttp.sslverifyfalse+}>expect&&+gitconfig--get-urlmatchHTTPhttps://good.example.com>actual&&+test_cmpexpectactual&&++echohttp.sslverify>expect&&+gitconfig--get-urlmatchHTTPhttps://more.example.com.au>actual&&test_cmpexpectactual'
@@ -63,6 +63,49 @@ static int append_normalized_escapes(struct strbuf *buf,return1;}+staticconstchar*end_of_token(constchar*s,intc,size_tn)+{+constchar*next=memchr(s,c,n);+if(!next)+next=s+n;+returnnext;+}++staticintmatch_host(conststructurl_info*url_info,+conststructurl_info*pattern_info)+{+constchar*url=url_info->url+url_info->host_off;+constchar*pat=pattern_info->url+pattern_info->host_off;+inturl_len=url_info->host_len;+intpat_len=pattern_info->host_len;++while(url_len&&pat_len){+constchar*url_next=end_of_token(url,'.',url_len);+constchar*pat_next=end_of_token(pat,'.',pat_len);++if(pat_next==pat+1&&pat[0]=='*')+/* wildcard matches anything */+;+elseif((pat_next-pat)==(url_next-url)&&+!memcmp(url,pat,url_next-url))+/* the components are the same */+;+else+return0;/* found an unmatch */++if(url_next<url+url_len)+url_next++;+url_len-=url_next-url;+url=url_next;+if(pat_next<pat+pat_len)+pat_next++;+pat_len-=pat_next-pat;+pat=pat_next;+}++return(!url_len&&!pat_len);+}+staticchar*url_normalize_1(constchar*url,structurl_info*out_info,charallow_globs){/*
@@ -467,9 +510,7 @@ static int match_urls(const struct url_info *url,}/* check the host */-if(url_prefix->host_len!=url->host_len||-strncmp(url->url+url->host_off,-url_prefix->url+url_prefix->host_off,url->host_len))+if(!match_host(url,url_prefix))return0;/* host names do not match *//* check the port */
From: Patrick Steinhardt <hidden> Date: 2017-01-31 09:09:41
From: Patrick Steinhardt <redacted>
The `url_normalize` function is used to validate and normalize URLs. As
such, it does not allow for some special characters to be part of the
URLs that are to be normalized. As we want to allow using globs in some
configuration keys making use of URLs, namely `http.<url>.<key>`, but
still normalize them, we need to somehow enable some additional allowed
characters.
To do this without having to change all callers of `url_normalize`,
where most do not actually want globbing at all, we split off another
function `url_normalize_1`. This function accepts an additional
parameter `allow_globs`, which is subsequently called by `url_normalize`
with `allow_globs=0`.
As of now, this function is not used with globbing enabled. A caller
will be added in the following commit.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
From: Patrick Steinhardt <hidden> Date: 2017-01-31 09:09:50
Hi,
this is version 5 of my patch series. The previous version can
be found at [1]. The use case is to be able to configure an HTTP
proxy for all subdomains of a domain where there are hundreds of
subdomains.
This includes only a single change, interdiff is included below.
The previous version had an embarassing bug because a variable
was not properly initialized in all cases, leading to undefined
behavior. I also verified that the patches work on top of
4e59582ff (Seventh batch for 2.12, 2017-01-23), where Junio
reported the test failures.
Regards
Patrick
Patrick Steinhardt (5):
mailmap: add Patrick Steinhardt's work address
urlmatch: enable normalization of URLs with globs
urlmatch: split host and port fields in `struct url_info`
urlmatch: include host in urlmatch ranking
urlmatch: allow globbing for the URL host part
.mailmap | 1 +
Documentation/config.txt | 5 +-
t/t1300-repo-config.sh | 105 ++++++++++++++++++++++++++++++++++++
urlmatch.c | 138 +++++++++++++++++++++++++++++++++++------------
urlmatch.h | 12 +++--
5 files changed, 220 insertions(+), 41 deletions(-)
--
2.11.0
From: Patrick Steinhardt <hidden> Date: 2017-01-31 09:09:52
From: Patrick Steinhardt <redacted>
The `url_info` structure contains information about a normalized URL
with the URL's components being represented by different fields. The
host and port part though are to be accessed by the same `host` field,
so that getting the host and/or port separately becomes more involved
than really necessary.
To make the port more readily accessible, split up the host and port
fields. Namely, the `host_len` will not include the port length anymore
and a new `port_off` field has been added which includes the offset to
the port, if available.
The only user of these fields is `url_normalize_1`. This change makes it
easier later on to treat host and port differently when introducing
globs for domains.
Signed-off-by: Patrick Steinhardt <redacted>
---
urlmatch.c | 16 ++++++++++++----
urlmatch.h | 9 +++++----
2 files changed, 17 insertions(+), 8 deletions(-)
@@ -464,11 +466,17 @@ static int match_urls(const struct url_info *url,usermatched=1;}-/* check the host and port */+/* check the host */if(url_prefix->host_len!=url->host_len||strncmp(url->url+url->host_off,url_prefix->url+url_prefix->host_off,url->host_len))-return0;/* host names and/or ports do not match */+return0;/* host names do not match */++/* check the port */+if(url_prefix->port_len!=url->port_len||+strncmp(url->url+url->port_off,+url_prefix->url+url_prefix->port_off,url->port_len))+return0;/* ports do not match *//* check the path */pathmatchlen=url_match_prefix(
@@ -18,11 +18,12 @@ struct url_info {size_tpasswd_len;/* length of passwd; if passwd_off != 0 butpasswd_len==0,anemptypasswdwasgiven*/size_thost_off;/* offset into url to start of host name (0 => none) */-size_thost_len;/* length of host name; this INCLUDES any ':portnum';+size_thost_len;/* length of host name;*fileurlsmayhavehost_len==0*/-size_tport_len;/* if a portnum is present (port_len != 0), it has-*thislength(excludingtheleading':')atthe-*endofthehostname(always0forfileurls)*/+size_tport_off;/* offset into url to start of port number (0 => none) */+size_tport_len;/* if a portnum is present (port_off != 0), it has+*thislength(excludingtheleading':')starting+*fromport_off(always0forfileurls)*/size_tpath_off;/* offset into url to the start of the url path;*thiswillalwayspointtoa'/'character*aftertheurlhasbeennormalized*/
From: Patrick Steinhardt <hidden> Date: 2017-01-31 09:09:55
From: Patrick Steinhardt <redacted>
In order to be able to rank positive matches by `urlmatch`, we inspect
the path length and user part to decide whether a match is better than
another match. As all other parts are matched exactly between both URLs,
this is the correct thing to do right now.
In the future, though, we want to introduce wild cards for the domain
part. When doing this, it does not make sense anymore to only compare
the path lengths. Instead, we also want to compare the domain lengths to
determine which of both URLs matches the host part more closely.
Signed-off-by: Patrick Steinhardt <redacted>
---
t/t1300-repo-config.sh | 33 ++++++++++++++++++++++++++++
urlmatch.c | 59 +++++++++++++++++++++++++++++---------------------
urlmatch.h | 3 ++-
3 files changed, 69 insertions(+), 26 deletions(-)
@@ -1177,6 +1177,39 @@ test_expect_success 'urlmatch' 'test_cmpexpectactual'+test_expect_success'urlmatch favors more specific URLs''+cat>.git/config<<-\EOF&&+[http"https://example.com/"]+cookieFile=/tmp/root.txt+[http"https://example.com/subdirectory"]+cookieFile=/tmp/subdirectory.txt+[http"https://user@example.com/"]+cookieFile=/tmp/user.txt+[http"https://averylonguser@example.com/"]+cookieFile=/tmp/averylonguser.txt+EOF++echohttp.cookiefile/tmp/root.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://example.com/subdirectory/nested>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/user.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://user@example.com/>actual&&+test_cmpexpectactual&&++echohttp.cookiefile/tmp/subdirectory.txt>expect&&+gitconfig--get-urlmatchHTTPhttps://averylonguser@example.com/subdirectory>actual&&+test_cmpexpectactual+'+# good section hygiene test_expect_failure'unsetting the last key in a section removes header''cat>.git/config<<-\EOF&&
@@ -445,8 +445,8 @@ static int match_urls(const struct url_info *url,*containedausernameorfalseifurl_prefixdidnothavea*username.Ifthereisnomatch*exactusermatchisleftuntouched.*/-intusermatched=0;-intpathmatchlen;+charusermatched=0;+size_tpathmatchlen;if(!url||!url_prefix||!url->url||!url_prefix->url)return0;
@@ -483,22 +483,38 @@ static int match_urls(const struct url_info *url,url->url+url->path_off,url_prefix->url+url_prefix->path_off,url_prefix->url_len-url_prefix->path_off);+if(!pathmatchlen)+return0;/* paths do not match */-if(pathmatchlen&&exactusermatch)-*exactusermatch=usermatched;-returnpathmatchlen;+if(match){+match->hostmatch_len=url_prefix->host_len;+match->pathmatch_len=pathmatchlen;+match->user_matched=usermatched;+}++return1;+}++staticintcmp_matches(conststructurlmatch_item*a,+conststructurlmatch_item*b)+{+if(a->hostmatch_len!=b->hostmatch_len)+returna->hostmatch_len<b->hostmatch_len?-1:1;+if(a->pathmatch_len!=b->pathmatch_len)+returna->pathmatch_len<b->pathmatch_len?-1:1;+if(a->user_matched!=b->user_matched)+returnb->user_matched?-1:1;+return0;}inturlmatch_config_entry(constchar*var,constchar*value,void*cb){structstring_list_item*item;structurlmatch_config*collect=cb;-structurlmatch_item*matched;+structurlmatch_itemmatched={0};structurl_info*url=&collect->url;constchar*key,*dot;structstrbufsynthkey=STRBUF_INIT;-size_tmatched_len=0;-intuser_matched=0;intretval;if(!skip_prefix(var,collect->section,&key)||*(key++)!='.'){
@@ -528,24 +544,17 @@ int urlmatch_config_entry(const char *var, const char *value, void *cb)item=string_list_insert(&collect->vars,key);if(!item->util){-matched=xcalloc(1,sizeof(*matched));-item->util=matched;+item->util=xcalloc(1,sizeof(matched));}else{-matched=item->util;-/*-*Isourmatchshorter?Isourmatchthesame-*length,andwithoutuserwhilethecurrent-*candidateiswithuser?Thenwecannotuseit.-*/-if(matched_len<matched->matched_len||-((matched_len==matched->matched_len)&&-(!user_matched&&matched->user_matched)))+if(cmp_matches(&matched,item->util)<=0)+/*+*Ourmatchisworsethantheoldone,+*wecannotuseit.+*/return0;-/* Otherwise, replace it with this one. */}-matched->matched_len=matched_len;-matched->user_matched=user_matched;+memcpy(item->util,&matched,sizeof(matched));strbuf_addstr(&synthkey,collect->section);strbuf_addch(&synthkey,'.');strbuf_addstr(&synthkey,key);