lychee

mirror of https://github.com/Hopiu/lychee.git synced 2026-05-16 17:51:08 +00:00

Author	SHA1	Message	Date
Matthias Endler	da46734c54	Extend response stats in verbose mode (#882 )	2022-12-20 10:43:01 +01:00
Matthias Endler	6df1c378ec	Fix Rust 1.66 clippy lints (#879 )	2022-12-19 14:28:10 +01:00
Matthias	7d435f2155	Add more markdown extensions (#866 )	2022-12-12 18:26:42 +01:00
Matthias	ef391cea50	Recursively skip verbatim elements (#847 )	2022-12-12 01:06:45 +01:00
Matthias	9eeea250cd	Exclude <script> tags by default (#848 ) This is a naive approach to exclude script tags from getting checked. The reason is that the tag leads to a lot of false-positives (e.g. `//unpkg.com/docsify-edit-on-github@1` within a script block gets detected as an e-mail address). A more thorough approach would be the use of a tree-builder in html5gum and html5ever, but this could have a negative performance impact. I also did not want to add a new flag (e.g. `--include-scripts`) for this setting because the current set of flags around exclusion/inclusion is already quite long. Fixes #821.	2022-11-29 00:38:43 +01:00
Matthias	982d978e47	Add different verbosity levels (#824 ) More granular verbosity levels have been asked for repeatedly. To enable that we're moving to [env_logger] and [clap-verbosity-flag] to provide more flexible verbosity settings. Also tackles #661, #709 Lays the groundwork for tackling #268 https://github.com/rust-cli/env_logger https://github.com/clap-rs/clap-verbosity-flag	2022-11-28 23:25:33 +01:00
Matthias	b479a5810e	Allow overriding accepted status codes for cached URIs (#843 ) Fixes #840	2022-11-28 12:23:07 +01:00
Matthias	765f7adb12	Don't check example mail addresses by default (#815 ) This was an oversight so far that became apparent after our recent fix for email addreses with query params (e.g. `test@example.com?subject=test`). The parsing of email addresses has improved and so we detect more mail addresses, but we didn't check if they belonged to an example domain, causing false-positive checks.	2022-11-08 23:46:32 +01:00
Matthias	d61105edbb	Fix parsing error of email addresses with query params (#809 ) Email addresses with query parameters often get used in contact forms on websites. They can also be found in other documents like Markdown. A common use-case is to add a subject line to the email as a parameter e.g. `mailto:mail@example.com?subject="Hello"`. Previously we handled such cases incorrectly by recognizing them as files. The reason was that our email parsing was too strict to allow for that use-case. With `email_address` we switched to a more permissive parser. Note that this does not affect the actual address email checking, as this is still done `check-if-email-exists`, which has more strict check functionality.	2022-11-05 23:40:33 +01:00
Matthias	94dda21326	Fix clippy lints	2022-09-27 18:17:37 +02:00
dependabot[bot]	226546091b	Bump check-if-email-exists from 0.8.31 to 0.9.0 (#735 ) * Bump check-if-email-exists from 0.8.31 to 0.9.0 Bumps [check-if-email-exists](https://github.com/reacherhq/check-if-email-exists) from 0.8.31 to 0.9.0. - [Release notes](https://github.com/reacherhq/check-if-email-exists/releases) - [Changelog](https://github.com/reacherhq/check-if-email-exists/blob/master/CHANGELOG.md) - [Commits](https://github.com/reacherhq/check-if-email-exists/compare/v0.8.31...v0.9.0) --- updated-dependencies: - dependency-name: check-if-email-exists dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Update usage Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Matthias <matthias-endler@gmx.net>	2022-08-16 12:35:34 +02:00
Matthias	6a49cedc16	Check Twitter URLs using nitter.net (#731 ) Use an alternative Twitter frontend, which works more reliably than using Twitter directly.	2022-08-12 22:46:35 +02:00
Matthias	69f387c1bd	Markdown-status (#729 ) * Fix typos * Add status code description to markdown output	2022-08-11 22:08:05 +02:00
Walter Beller-Morales	6d40a2ab7b	Update to gracefully handle nonexistent relative paths (#691 ) * Update Input::new to gracefully handle nonexistent relative paths * Add test checking Input::new can handle real relative paths * Add better pre-conditions to Input::new tests * Add integration tests for handling relative paths in lychee-bin * Update lychee-lib/src/types/input.rs	2022-07-22 17:15:55 +02:00
Matthias	6fae93f2da	Skip caching unsupported and excluded URLs (#692 ) As discussed in https://github.com/lycheeverse/lychee/issues/647#issuecomment-1170773449, it does not make much sense to cache unsupported and excluded URLs. Unsupported URLs might be supported in the future and caching them would mean they won't get checked then. Excluded URLs were excluded for a reason and should not appear in the cache. Furthermore they might not be excluded in a consecutive run, leading to a false-positive.	2022-07-17 18:40:45 +02:00
Walter Beller-Morales	9ad53f97a2	Fix deserialize of lycheecache status codes (#685 ) * Add custom deserializer for `CacheStatus` to properly classify status codes * Add CLI integration tests to check .lycheecache behavior * Add comment to explain conflict between cache and accept flags	2022-07-15 22:45:24 +02:00
Matthias	fb367ef43a	Add http://www.w3.org/1999/xlink to list of false positives (#664 )	2022-07-01 12:11:57 +02:00
Matthias	487d88cefe	Add test for mailto address with query params (#655 )	2022-06-29 10:19:17 +02:00
dependabot[bot]	231939af82	Bump html5gum from 0.4.0 to 0.5.1 (#658 ) * Bump html5gum from 0.4.0 to 0.5.1 Bumps [html5gum](https://github.com/untitaker/html5gum) from 0.4.0 to 0.5.1. - [Release notes](https://github.com/untitaker/html5gum/releases) - [Commits](https://github.com/untitaker/html5gum/compare/0.4.0...0.5.1) --- updated-dependencies: - dependency-name: html5gum dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Update html5gum Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Matthias <matthias-endler@gmx.net>	2022-06-23 00:07:28 +02:00
Markus Unterwaditzer	f1ae22da09	Replace lazy hashset with matches! (#656 ) * Replace lazy hashset with matches! llvm will typically create much faster code than accessing a hashset at runtime source: trust me bro * cargo fix * cargo fmt * shorten docstring	2022-06-18 19:00:07 +02:00
Matthias	84de43c554	Refactor request types (#637 )	2022-06-03 20:13:07 +02:00
Matthias	9b4dfadffd	Fix parsing errors with config options (#632 )	2022-05-31 19:43:46 +02:00
Matthias	f33b897d5d	Exclude example domains as per RFC 2606 from checking (#627 ) Unfortunately it's not possible to automatically enable features for `cargo test`. See https://github.com/rust-lang/cargo/issues/2911. As a workaround to allow for using example domains for unit- and integration tests, we introduce a new feature, `check_example_domains`, which is disabled by default for normal users. The feature gets activated for the integration test which checks that the example domain exclusion works as expected.	2022-05-29 21:42:00 +02:00
Matthias	22fecfc056	Add support for URI remapping (#620 ) Remaps allow mapping from a URI pattern to a different URI. The syntax is ``` lychee --remap 'https://example.com http://127.0.0.1' ``` Some use-cases are - Testing URIs prior to production deployment - Testing URIs behind a proxy Be careful when using this feature because checking every link against a large set of regular expressions has a performance impact. Also there are no constraints on the URI mapping, so the rules might contradict with each other. Remap rules get applied in order of definition to every input URI.	2022-05-29 21:41:22 +02:00
Matthias	363b95fe5f	Add support for excluding paths from link checking (#623 ) This change deprecates `--exclude-file` as it was ambiguous. Instead, `--exclude-path` was introduced to support excluding paths to files and directories that should not be checked. Furthermore, `.lycheeignore` is now the only way to exclude URL patterns.	2022-05-29 17:27:09 +02:00
Matthias	571b49410c	Extend reqwest client settings (#617 ) This sets a HTTP connect timeout (for stability) and a TCP keepalive (for performance). The connect timeout should help with flaky servers, which would block the runtime and therefore other requests. The keepalive helps when making many requests to the same host. This is a very common pattern for checking internal documentation, which is an important use-case of lychee. The settings are currently not configurable by the user and set to sane defaults. We might make this configurable in the future if there is demand to do so.	2022-05-13 18:51:11 +02:00
Matthias	8c0a32d81d	Refactor response formatting (#599 ) * Add support for raw formatter (no color) * Introduce ResponseFormatter trait * Pass the same params to every cli command * Update dependencies * Remove pretty_assertions dependency (latest version doesn't build)	2022-04-25 19:19:36 +02:00
Matthias	a607b853c9	Move to downstream optimization for short strings (#600 ) Skipping to parse very short strings was merged into linkify so our own workaround is unnecessary https://github.com/robinst/linkify/pull/34	2022-04-25 19:18:50 +02:00
Matthias	da7bbf113d	Remove unnecessary Ok wrapper	2022-04-12 01:39:38 +02:00
Matthias	6ebc9fed4b	Reset nofollow in html5gum start tag (#584 )	2022-04-06 00:49:00 +02:00
Matthias	debe958766	Add support for nofollow (#572 )	2022-04-04 10:32:00 +02:00
Matthias	03d28820bb	Extract more status information from reqwest (#577 ) Recently we cleaned up the commandline output to trim away redundant information like the URL, which occured twice. Unfortunately we also removed helpful information from reqwest, which could support the user in troubleshooting unexpected errors. This commit reverts that. We now extract the meaningful information from reqwest, without being too verbose. For that we have to depend on the string output for the reqwest error, but it's better than hiding that information from the user. It is fragile as it depends on the reqwest internals, but in the worst case we simply return the full error text in case our parsing won't work.	2022-04-02 14:37:03 +02:00
Matthias	5ad7b14bdd	Regression: Ignore invalid URLs (#571 ) With the refactoring the URL checking as a workaround for the upstream reqwest panic on invalid URLs, we introduced a regression, which caused unsupported URL schemes to show up as errors in the lychee output. This commit changes the behavior such that invalid schemes get ignored again by making a differentiation between truly invalid URIs which would make reqwest panic, and ones which are valid but just not handled by reqwest. The check was moved to `check_website` such that the invalid URIs would not be checked three times in a loop before erroring out.	2022-03-27 23:22:46 +02:00
Matthias	36d3195c68	Cache verbosity issue (fixes #562 )	2022-03-27 14:48:09 +02:00
Matthias	743d386252	Allow input URLs without scheme (fixes #567 ) This requires `Input::new` to return a `Result`, because the URL parsing could fail when prepending `http://`. We use http instead of https, because curl does as well: `70ac27604a/lib/urlapi.c (L1104-L1124)` Missing files will be interpreted as URLs from the command line and these can be invalid, but that's not seen as an error anymore.	2022-03-27 01:27:27 +01:00
Matthias	d616177a99	Implement excluding code blocks (#523 ) This is done in the extractor to avoid unnecessary allocations.	2022-03-26 10:42:56 +01:00
Matthias	77b1724881	Optimize plaintext extractor for small strings (#565 ) Immediately return for very small strings which cannot be valid URIs. The shortest valid URI without a scheme might be g.cn (Google China) At least I am not aware of a shorter one. We set this as a lower threshold for parsing URIs from plaintext to avoid false-positives and as a slight performance optimization, which could add up for big files. This threshold might be adjusted in the future.	2022-03-23 23:06:49 +01:00
Matthias	e1d112dbab	Remove `missing_panic_doc` (#561 )	2022-03-22 21:02:56 +01:00
Matthias	45de5c763e	Avoid reqwest panic on invalid URIs (#557 )	2022-03-22 13:15:11 +01:00
Matthias	ceb185e579	Add more comments to path methods (#543 )	2022-03-08 13:50:54 +01:00
Matthias	8097bfa408	Print Github token error once at the end (#537 ) Print original reqwest error for every Github link. It contains more information about the underlying error. Only print a message about the Github token at the end if it's not set and there were Github errors.	2022-03-03 10:04:55 +01:00
Matthias	4c51fce22f	Fix broken pipe error on failing writes to stdout (#535 ) Make sure that broken pipes (e.g. when a reader of a pipe prematurely exits during execution) get handled gracefully. This change also moves some error messages to stderr by using eprintln. More info: https://github.com/jez/as-tree/issues/15	2022-03-02 23:39:54 +01:00
Matthias	0fc5fc9ffe	Print errors with a different format for easier clickability (fixes #532 )	2022-03-01 16:58:04 +01:00
Matthias	05bd3817ee	Make retry wait time configurable (#525 )	2022-02-24 12:24:57 +01:00
Matthias	41b291037a	Response output overhaul (#524 ) Clean up the response output. Superfluous information was removed and the formatting was changed to make the output more readable to humans.	2022-02-23 17:28:14 +01:00
Lucius Hu	70ebe45117	Improved IPv6 filtering support (#501 ) This commit uses crate `ip_network` to determine whether an IPv6 address is link-local or unique local. Note that this extra dependencies can be removed once rust-lang/rust#27709 is stabilized. Co-authored-by: Lucius Hu <lebensterben@users.noreply.github.com> Co-authored-by: Matthias <matthias-endler@gmx.net>	2022-02-22 10:39:44 +01:00
Matthias	ba276cd51b	Error cleanup (#510 ) * Add more fine-grained error types; remove generic IO error * Update error message for missing file * Remove missing `Error` suffix * Rename ErrorKind::Github to ErrorKind::GithubRequest for consistency with NetworkRequest	2022-02-19 01:44:00 +01:00
Matthias	812663d832	Prevent flaky tests (#514 ) Move from example.org to example.com, which seems to be more permissive for testing	2022-02-18 10:29:49 +01:00
Lucius Hu	6d56c6b55c	Replace plain String with SecretString for GitHub token (#509 ) This commit changed the type of `lychee-lib::ClientBuilder::github_token` from `String` to `secrecy::SecretString` to fortify the secret management within our program. Note that this won't affect TOML configuration of `lychee-bin` because `serde::Deserialize` is still implemented for `SecretString`.	2022-02-13 13:53:46 +01:00
Matthias	47df7780fe	Use captured identifiers in format strings (#507 ) Makes for arguably cleaner-looking code. The downside is that the MSRV is 1.58 https://blog.rust-lang.org/2022/01/13/Rust-1.58.0.html Given that nobody uses lychee as a library yet and we have precompiled binaries, it's an acceptable tradeoff. My little research revealed that this is a much-liked feature: https://twitter.com/matthiasendler/status/1483895557621960715	2022-02-12 10:51:52 +01:00
Lucius Hu	53c41b03d8	replace hubcaps by octocrab (#502 ) This commit replaced `hubcaps` by `octocrab`, which has more downloads per month and receives more frequent release updates. The caveats are: 1. When instantiating the API client, `octocrab` doesn't offer you a way to specify custom user-agent. But I would argue that, at least presently, this doesn't seem to cause issues. 2. `octocrab` doesn't export as much details of its error types as `hubcaps` does. So we will have fewer control on the display of the error message. But I would also argue that this is not really important. Though we should do more tests to make sure the error looks good enough. * hide implementation details in error message Co-authored-by: Lucius Hu <lebensterben@users.noreply.github.com>	2022-02-11 23:43:47 +01:00
Lucius Hu	476a048350	lychee-lib::client reworked (#500 ) This commit mainly added or improved documentation for `lychee-lib::client` module. But it also contains a few API changes: - `ClientBuilder::client()` now consumes itself instead of taking a reference. This helps to avoid a few unnecessary clones. - `ClientBuilder::build_filter()` was a private function and is inlined to avoid unnecessary clones. - Added a new crate-scoped function `Uri::set_scheme()`. * added notes on deprecated site-local network Co-authored-by: Lucius Hu <lebensterben@users.noreply.github.com>	2022-02-10 00:04:48 +01:00
Markus Unterwaditzer	68d09f7e5b	Add html5gum as alternative link extractor (#480 ) html5gum is a HTML parser that offers lower-level control over which tokens actually get created and are tracked. As such, the extractor doesn't allocate anything tokens it doesn't care about. On some benchmarks it provides a substantial performance boost. The old parser, html5ever is still available by setting the `LYCHEE_USE_HTML5EVER=1` env var.	2022-02-07 22:54:47 +01:00
Matthias	6635863746	Add Alpine page for benchmark; refactor code (#481 )	2022-01-27 23:42:06 +01:00
Matthias	97b06230fc	Add missing Github exclusions; sort entries (#473 )	2022-01-21 23:54:59 +01:00
Matthias	5802ae912c	Fix bugs in extractor; reduce allocs (#464 ) When URLs couldn't be extracted from a tag, we ran a plaintext search, but never added the newly found urls to the vec of extracted urls. Also tried to make the code a little more idiomatic	2022-01-16 02:13:38 +01:00
Matthias	6e757fa20e	Add more information about mail errors (#463 )	2022-01-14 22:22:53 +01:00
Matthias	994aadf6a1	Simplify error messages (#462 ) Using pattern matching to make the hubcaps and reqwest error messages a little shorter and (subjectively) more readable.	2022-01-14 15:26:13 +01:00
Matthias	ac490f9c53	Add caching functionality (v2) (#443 ) A while ago, caching was removed due to some issues (see #349). This is a new implementation with the following improvements: * Architecture: The new implementation is decoupled from the collector, which was a major issue in the last version. Now the collector has a single responsibility: collecting links. This also avoids race-conditions when running multiple collect_links instances, which probably was an issue before. * Performance: Uses DashMap under the hood, which was noticeably faster than Mutex<HashMap> in my tests. * Simplicity: The cache format is a CSV file with two columns: URI and status. I decided to create a new struct called CacheStatus for serialization, because trying to serialize the error kinds in Status turned out to be a bit of a nightmare and at this point I don't think it's worth the pain (and probably isn't idiomatic either). This is an optional feature. Caching only gets used if the `--cache` flag is set.	2022-01-14 15:25:51 +01:00
Matthias	1e76e82811	Add test for nonexistent Github file	2022-01-12 09:25:12 +01:00
Matthias	48c8153e11	Refactor Github checking; add docs	2022-01-12 09:25:12 +01:00
Matthias	50d7b05736	Conditionally compile constructors for GithubUri for tests	2022-01-12 09:25:12 +01:00
Matthias	8d445a3a4b	Be more permissive around private GH repos The Github API doesn't handle checking individual files inside repos or paths like `github.com/org/repo/issues`, so we are more permissive and only check for repo existence. This is the only way to get a basic check for private repos. Public repos are not affected and should work with a normal check.	2022-01-12 09:25:12 +01:00
Matthias	e91c0c60f0	Only accept two path segments (org/repo) for Github API check	2022-01-12 09:25:12 +01:00
Matthias	7667842bb6	Strip `.git` suffix from Github URLs (#384 )	2022-01-12 09:25:12 +01:00
Matthias	21f3160b71	Make retries configurable; align constants (#446 ) Using the same default values for the library and the binary now but tweaked the values a bit for slightly faster performance.	2022-01-07 01:03:10 +01:00
Matthias	388bbbe7b0	Exclude known false-positives from Github API check (#445 ) Fixes https://github.com/lycheeverse/lychee/issues/431	2022-01-06 00:33:53 +01:00
Matthias	dd48466d9a	Add missing test for local links in plaintext files (#444 )	2022-01-05 12:51:14 +01:00
Matthias	01393b34a2	Upgrade to Rust 2021 (#427 )	2021-12-17 01:32:13 +01:00
Matthias	83182c29ca	Fix JSON serialization (#426 ) We recently removed the custom serialization for InputSource. This causes the JSON formatter to fail with "key must be a string". Add it back and add a comment on why this is needed.	2021-12-16 23:55:04 +01:00
Matthias	166c86c30e	Use tokenizer for extraction; add benchmark (#424 ) This avoids creating a DOM tree for link extraction and instead uses a `TokenSink` for on-the-fly extraction. In hyperfine benchmarks it was about 10-25% faster than the master. Old: 4.557 s ± 0.404 s New: 3.832 s ± 0.131 s The performance fluctuates a little less as well. Some missing element/attribute pairs were also added, which contain links according to the HTML spec. These occur very rarely, but it's good to parse them for completeness' sake. Furthermore tried to clean up a lot of papercuts around our types. We now differentiate between a `RawUri` (stringy-types) and a Uri, which is a properly parsed `URI` type. The extractor now only deals with extracting `RawUri`s while the collector creates the request objects.	2021-12-16 18:45:52 +01:00
Matthias	c41ba64a69	Max concurrency moved to check (#419 ) Concurrency is defined by the channel size consuming from the request stream in `check`	2021-12-07 11:52:40 +01:00
Matthias	3d5135668b	Improve concurrency with streams (#330 ) * Move to from vec to streams Previously we collected all inputs in one vector before checking the links, which is not ideal. Especially when reading many inputs (e.g. by using a glob pattern), this could cause issues like running out of file handles. By moving to streams we avoid that scenario. This is also the first step towards improving performance for many inputs. To stay as close to the pre-stream behaviour, we want to stop processing as soon as an Err value appears in the stream. This is easiest when the stream is consumed in the main thread. Previously, the stream was consumed in a tokio task and the main thread waited for responses. Now, a tokio task waits for responses (and displays them/registers response stats) and the main thread sends links to the ClientPool. To ensure that the main thread waits for all responses to have arrived before finishing the ProgressBar and printing the stats, it waits for the show_results_task to finish. * Return collected links as Stream * Initialize ProgressBar without length because we can't know the amount of links without blocking * Handle stream results in main thread, not in task * Add basic directory support using jwalk * Add test for HTTP protocol file type (http://) * Remove deadpool (once again): Replaced with `futures::StreamExt::for_each_concurrent`. * Refactor main; fix tests * Move commands into separate submodule * Simplify input handling * Simplify collector * Remove unnecessary unwrap * Simplify main * cleanup check * clean up dump command * Handle requests in parallel * Fix formatting and lints Co-authored-by: Timo Freiberg <self@timofreiberg.com>	2021-12-01 18:25:11 +01:00
Matthias	d96c1269ff	Use thiserror for error handling (#399 ) This removes some boilerplate and is arguably better than handwriting the error handling code for maintainability and avoid inconsitent functionality for the error variants. thiserror is also the de-facto standard for library error types as of today.	2021-11-20 01:42:50 +01:00
Matthias	b97fda34d0	Add support for different output formats (compact, detailed, markdown) (#375 )	2021-11-18 00:44:48 +01:00
Markus Unterwaditzer	d3ed133f10	Remove srcset attribute from list of "link" attrs (#393 ) * Remove srcset attribute from list of "link" attrs Fix #390 * Add test for srcset * Add note about srcSet links * add real support for srcset Co-authored-by: Matthias <matthias-endler@gmx.net>	2021-11-16 22:58:10 +01:00
Matthias	69e5d56687	Add more known false positive schema domains (#376 ) See https://github.com/lycheeverse/lychee-action/issues/53	2021-10-31 14:53:40 +01:00
dependabot[bot]	d3a72d3816	Bump deadpool from 0.7.0 to 0.9.1 (#371 ) * Bump deadpool from 0.7.0 to 0.9.1 Bumps [deadpool](https://github.com/bikeshedder/deadpool) from 0.7.0 to 0.9.1. - [Release notes](https://github.com/bikeshedder/deadpool/releases) - [Changelog](https://github.com/bikeshedder/deadpool/blob/master/CHANGELOG.md) - [Commits](https://github.com/bikeshedder/deadpool/compare/deadpool-v0.7.0...deadpool-v0.9.1) --- updated-dependencies: - dependency-name: deadpool dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Attempt fix for deadpool v0.8.0+ (#372) Signed-off-by: MichaIng <micha@dietpi.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: MichaIng <micha@dietpi.com>	2021-10-28 02:05:58 +02:00
Matthias	47426c6971	Fix typos, grammar	2021-10-28 02:05:35 +02:00
MichaIng	0870f0bc9e	Add http://www.w3.org/2000/svg to known false positives (#359 ) It has no forced HTTPS rewrite, but sets the HSTS header. Access otherwise works fine, so similar to http://www.w3.org/1999/xhtml it is basically to avoid lychee failures when --require-https was defined. Signed-off-by: MichaIng <micha@dietpi.com>	2021-10-11 00:40:27 +02:00
Jorge Luis Betancourt	174331d983	Extract base from the source URL if `--base` is empty (#358 ) When running lychee against a remote URL all relative links are ignored by default because `--base` is normally not set. A good default in this case is to automatically use the base domain from the source URL. Setting `--base` overrides the automatic source extraction from the source URL (same behaviour as we currently have).	2021-10-10 02:42:01 +02:00
Matthias	dd9e24b7f4	support uppercase filenames; add tests	2021-10-09 22:20:22 +02:00
Matthias	175342baf4	Merge branch 'master' of github.com:lycheeverse/lychee	2021-10-09 21:17:41 +02:00
Matthias	bdcd6f87bf	Make error message for broken file links more understandable	2021-10-09 21:17:37 +02:00
Matthias	56726f41fc	Add back connection pool (#355 )	2021-10-08 13:08:44 +02:00
MichaIng	961f12e58e	Remove cache from collector and remove custom reqwest client pool * Reqwest comes with its own request pool, so there's no need in adding another layer of indirection. This also gets rid of a lot of allocs. * Remove cache from collector * Improve error handling and documentation * Add back test for request caching in single file Signed-off-by: MichaIng <micha@dietpi.com> Co-authored-by: Matthias <matthias-endler@gmx.net>	2021-10-07 18:07:18 +02:00
Matthias	a7f809612d	Refactor extractor (#354 ) This avoids sending URLs back and forth between the different parsers. Also, it should allow for future optimizations to reduce allocs.	2021-10-07 12:51:02 +02:00
MichaIng	b648b5e914	Imply "localhost" when loopback IPs are excluded (#351 ) as "localhost" is usually mapped via "hosts" file to a loopback IP address. Resolves: https://github.com/lycheeverse/lychee/issues/319 Signed-off-by: MichaIng <micha@dietpi.com>	2021-10-06 11:33:23 +02:00
Matthias	251332efe2	Cache `absolute_path` to decrease allocations (#346 ) * Cache `absolute_path` to decrease allocations While profiling local file handling, I noticed that resolving paths was taking a significant amount of time. It also caused quite a few allocations. By caching the path and using a constant value for the current directory, we can reduce the number of allocs by quite a lot. For example, when testing on the sentry documentation, we do 50,4% less allocations in total now. That's just a single test-case of course, but it's probably also helping in many other cases as well. * Defer to_string for attr.value to reduce allocs * Use Tendrils instead of Strings for parsing (another ~1.5% less allocs) * Move option parsing code into separate module * Handle base dir more correctly * Temporarily disable dry run	2021-10-05 01:37:43 +02:00
Matthias	3b41c4c375	Silently ignore absolute paths without base (fixes #320 ) (#338 )	2021-09-20 11:13:30 +02:00
Matthias	21ea0fd033	Add support for tokio-console (#318 ) This allows troubleshooting and improving async Rust code. It is an optional feature that is still experimental (but can be quite helpful)	2021-09-12 18:10:23 +02:00
Matthias	de55fbd178	Add TODO for fixing URL encoding for paths	2021-09-09 19:31:49 +02:00
Matthias	d7436575eb	formatting	2021-09-09 14:43:40 +02:00
Matthias	2a4170eade	Add test for `+` encoding	2021-09-09 14:42:09 +02:00
Matthias	a1acf7b0d0	Reintegrate master	2021-09-09 01:49:25 +02:00
Matthias	93948d7367	Avoid double-encoding already encoded destination paths E.g. `web%20site` becomes `web site`. That's because Url::from_file_path will encode the full URL in the end. This behavior cannot be configured. See https://github.com/lycheeverse/lychee/pull/262#issuecomment-915245411	2021-09-09 01:44:10 +02:00
Matthias	24ea2482d3	Update docs	2021-09-08 01:08:59 +02:00
Matthias	f3fe46a4d6	Merge branch 'master' of github.com:lycheeverse/lychee into local-files	2021-09-08 00:35:41 +02:00
Matthias	ffab0343fc	Revert refactor for removing params and fragments The refactored version was not equivalent. It could not handle fragments containing a question mark. See `67268ed598 (r703400238)`	2021-09-08 00:29:30 +02:00
Matthias	1246fa564c	Don't exlude mail on `exclude-all-private` (#316 )	2021-09-08 00:21:00 +02:00

1 2 3 4

186 commits