go-crawler, v1.9
Posted on

The library that implements crawling of all relative links for specified ones.
Supporting of a gzip compression of a sitemap.xml file.
Change Log
- crawling of all relative links for specified ones:
- extracting links from a
sitemap.xmlfile (optional):- supporting of a gzip compression of a
sitemap.xmlfile.
- supporting of a gzip compression of a
- extracting links from a
Features
- crawling of all relative links for specified ones:
- repeated extracting of relative links on error (optional):
- only specified repeat count;
- supporting of delay between repeats;
- delayed extracting of relative links (optional):
- reducing of a delay time by the time elapsed since the last request;
- using of individual delays for each thread;
- extracting links from a
sitemap.xmlfile (optional):- ignoring of the error on loading of the
sitemap.xmlfile:- logging of the received error;
- returning of an empty Sitemap instead;
- supporting of few
sitemap.xmlfiles for a single link:- processing of each
sitemap.xmlfile is done in a separate goroutine; - supporting of an outer generator for
sitemap.xmllinks:- generators:
- simple generator (it returns the
sitemap.xmlfile in the site root); - hierarchical generator (it returns the suitable
sitemap.xmlfile for each part of the URL path); - generator based on the
robots.txtfile;
- simple generator (it returns the
- supporting of grouping of generators:
- result of group generating is merged results of each generator in the group;
- generating concurrently:
- processing of each generator is done in a separate goroutine;
- generators:
- processing of each
- supporting of a Sitemap index file:
- supporting of a delay before loading of each
sitemap.xmlfile listed in the index;
- supporting of a delay before loading of each
- supporting of a gzip compression of a
sitemap.xmlfile;
- ignoring of the error on loading of the
- supporting of grouping of link extractors:
- result of group extracting is merged results of each extractor in the group;
- extracting links concurrently:
- processing of each link extractor is done in a separate goroutine;
- repeated extracting of relative links on error (optional):
- calling of an outer handler for an each found link:
- it's called directly during crawling;
- handling of links immediately after they have been extracted;
- passing of the source link in the outer handler;
- handling links filtered by a custom link filter (optional);
- handling links concurrently (optional);
- custom filtering of considered links:
- by relativity of a link (optional);
- by uniqueness of an extracted link (optional):
- supporting of sanitizing of a link before checking of uniqueness (optional);
- by a
robots.txtfile (optional):- customized user agent;
- supporting of grouping of link filters:
- result of group filtering is successful only when all filters are successful;
- parallelization possibilities:
- crawling of relative links in parallel;
- supporting of background working:
- automatic completion after processing all filtered links;
- simulate an unbounded channel of links to avoid a deadlock.
Examples
crawler.HandleLinksConcurrently() with processing a sitemap.xml file:
package main
import (
"compress/gzip"
"context"
"fmt"
"io"
stdlog "log"
"net/http"
"net/http/httptest"
"os"
"path"
"runtime"
"strings"
"sync"
"text/template"
"time"
"github.com/go-log/log/print"
"github.com/thewizardplusplus/go-crawler"
"github.com/thewizardplusplus/go-crawler/checkers"
"github.com/thewizardplusplus/go-crawler/extractors"
"github.com/thewizardplusplus/go-crawler/handlers"
"github.com/thewizardplusplus/go-crawler/models"
"github.com/thewizardplusplus/go-crawler/registers"
"github.com/thewizardplusplus/go-crawler/registers/sitemap"
"github.com/thewizardplusplus/go-crawler/sanitizing"
htmlselector "github.com/thewizardplusplus/go-html-selector"
)
type LinkHandler struct {
ServerURL string
}
func (handler LinkHandler) HandleLink(
ctx context.Context,
link models.SourcedLink,
) {
fmt.Printf(
"have got the link %q from the page %q\n",
handler.replaceServerURL(link.Link),
handler.replaceServerURL(link.SourceLink),
)
}
// replace the test server URL for reproducibility of the example
func (handler LinkHandler) replaceServerURL(link string) string {
return strings.Replace(link, handler.ServerURL, "http://example.com", -1)
}
// nolint: gocyclo
func RunServer() *httptest.Server {
return httptest.NewServer(http.HandlerFunc(func(
writer http.ResponseWriter,
request *http.Request,
) {
if request.URL.Path == "/robots.txt" {
sitemapLink :=
completeLinkWithHost("/sitemap_from_robots_txt.xml", request.Host)
fmt.Fprintf(writer, `
User-agent: go-crawler
Disallow: /2
Sitemap: %s
`, sitemapLink)
return
}
var links []string
switch request.URL.Path {
case "/sitemap.xml":
links = []string{"/1", "/2", "/hidden/1", "/hidden/2"}
case "/sitemap_from_robots_txt.xml":
links = []string{"/hidden/3", "/hidden/4"}
case "/hidden/1/sitemap.xml":
links = []string{"/hidden/5", "/hidden/6"}
case "/1/sitemap.xml", "/2/sitemap.xml", "/hidden/sitemap.xml":
links = []string{}
}
completeLinksWithHost(links, request.Host)
if links != nil {
writer.Header().Set("Content-Encoding", "gzip")
compressingWriter := gzip.NewWriter(writer)
defer compressingWriter.Close() // nolint: errcheck
// nolint: errcheck
renderTemplate(compressingWriter, links, `
<?xml version="1.0" encoding="UTF-8" ?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
{{ range $link := . }}
<url>
<loc>{{ $link }}</loc>
</url>
{{ end }}
</urlset>
`)
return
}
switch request.URL.Path {
case "/":
links = []string{"/1", "/2", "/2", "https://golang.org/"}
case "/1":
links = []string{"/1/1", "/1/2"}
case "/2":
links = []string{"/2/1", "/2/2"}
case "/hidden/1":
links = []string{"/hidden/1/test"}
}
completeLinksWithHost(links, request.Host)
// nolint: errcheck
renderTemplate(writer, links, `
<ul>
{{ range $link := . }}
<li>
<a href="{{ $link }}">{{ $link }}</a>
</li>
{{ end }}
</ul>
`)
}))
}
func completeLinkWithHost(link string, host string) string {
return "http://" + path.Join(host, link)
}
func completeLinksWithHost(links []string, host string) {
for index := range links {
if strings.HasPrefix(links[index], "/") {
links[index] = completeLinkWithHost(links[index], host)
}
}
}
// nolint: unparam
func renderTemplate(writer io.Writer, data interface{}, text string) error {
template, err := template.New("").Parse(text)
if err != nil {
return err
}
return template.Execute(writer, data)
}
func main() {
server := RunServer()
defer server.Close()
links := make(chan string, 1000)
links <- server.URL
var waiter sync.WaitGroup
waiter.Add(1)
logger := stdlog.New(os.Stderr, "", stdlog.LstdFlags|stdlog.Lmicroseconds)
// wrap the standard logger via the github.com/go-log/log package
wrappedLogger := print.New(logger)
crawler.HandleLinksConcurrently(
context.Background(),
runtime.NumCPU(),
links,
crawler.HandleLinkDependencies{
CrawlDependencies: crawler.CrawlDependencies{
LinkExtractor: extractors.RepeatingExtractor{
LinkExtractor: extractors.ExtractorGroup{
extractors.DefaultExtractor{
HTTPClient: http.DefaultClient,
Filters: htmlselector.OptimizeFilters(htmlselector.FilterGroup{
"a": {"href"},
}),
},
extractors.SitemapExtractor{
SitemapRegister: registers.NewSitemapRegister(
time.Second,
sitemap.GeneratorGroup{
sitemap.HierarchicalGenerator{
SanitizeLink: sanitizing.SanitizeLink,
},
sitemap.RobotsTXTGenerator{
RobotsTXTRegister: registers.NewRobotsTXTRegister(http.DefaultClient),
},
},
wrappedLogger,
sitemap.Loader{HTTPClient: http.DefaultClient}.LoadLink,
),
Logger: wrappedLogger,
},
},
RepeatCount: 5,
RepeatDelay: time.Second,
Logger: wrappedLogger,
SleepHandler: time.Sleep,
},
LinkChecker: checkers.CheckerGroup{
checkers.HostChecker{
Logger: wrappedLogger,
},
checkers.DuplicateChecker{
LinkRegister: registers.NewLinkRegister(
sanitizing.SanitizeLink,
wrappedLogger,
),
},
},
LinkHandler: handlers.CheckedHandler{
LinkChecker: checkers.DuplicateChecker{
// don't use here the link register from the duplicate checker above
LinkRegister: registers.NewLinkRegister(
sanitizing.SanitizeLink,
wrappedLogger,
),
},
LinkHandler: LinkHandler{
ServerURL: server.URL,
},
},
Logger: wrappedLogger,
},
Waiter: &waiter,
},
)
waiter.Wait()
// Unordered output:
// have got the link "http://example.com/1" from the page "http://example.com"
// have got the link "http://example.com/1/1" from the page "http://example.com/1"
// have got the link "http://example.com/1/2" from the page "http://example.com/1"
// have got the link "http://example.com/2" from the page "http://example.com"
// have got the link "http://example.com/2/1" from the page "http://example.com/2"
// have got the link "http://example.com/2/2" from the page "http://example.com/2"
// have got the link "http://example.com/hidden/1" from the page "http://example.com"
// have got the link "http://example.com/hidden/1/test" from the page "http://example.com/hidden/1"
// have got the link "http://example.com/hidden/2" from the page "http://example.com"
// have got the link "http://example.com/hidden/3" from the page "http://example.com"
// have got the link "http://example.com/hidden/4" from the page "http://example.com"
// have got the link "http://example.com/hidden/5" from the page "http://example.com/hidden/1/test"
// have got the link "http://example.com/hidden/6" from the page "http://example.com/hidden/1/test"
// have got the link "https://golang.org/" from the page "http://example.com"
}
Repository
Link: https://github.com/thewizardplusplus/go-crawler/tree/v1.9.
Content: code.
License: MIT.