Package org.apache.nutch.crawl
Class LinkDbMerger
java.lang.Object
org.apache.hadoop.conf.Configured
org.apache.nutch.crawl.LinkDbMerger
- All Implemented Interfaces:
Configurable,Tool
This tool merges several LinkDb-s into one, optionally filtering URLs through
the current URLFilters, to skip prohibited URLs and links.
It's possible to use this tool just for filtering - in that case only one LinkDb should be specified in arguments.
If more than one LinkDb contains information about the same URL, all inlinks
are accumulated, but only at most linkdb.max.inlinks inlinks will
ever be added.
If activated, URLFilters will be applied to both the target URLs and to any incoming link URL. If a target URL is prohibited, all inlinks to that target will be removed, including the target URL. If some of incoming links are prohibited, only they will be removed, and they won't count when checking the above-mentioned maximum limit.
- Author:
- Andrzej Bialecki
-
Nested Class Summary
Nested Classes -
Constructor Summary
Constructors -
Method Summary
Methods inherited from class org.apache.hadoop.conf.Configured
getConf, setConfMethods inherited from class java.lang.Object
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, waitMethods inherited from interface org.apache.hadoop.conf.Configurable
getConf, setConf
-
Constructor Details
-
LinkDbMerger
public LinkDbMerger() -
LinkDbMerger
-
-
Method Details
-
merge
- Throws:
Exception
-
createMergeJob
public static Job createMergeJob(Configuration config, Path linkDb, boolean normalize, boolean filter) throws IOException - Throws:
IOException
-
main
Run the job- Parameters:
args- input arguments for the job- Throws:
Exception- if there is an error running the job
-
run
-