ARCOMEM NEER -- Named Entity Evolution Recognizer 

---------------------------------

ARCOMEM NEER -- Named Entity Evolution Recognizer 
The ARCOMEM Named Entity Evolution Recognizer is a java module for 
named entity evolution tracking developed by the Leibniz Universitaet 
Hannover, L3S Research Center (http://www.L3S.de) as part of 
the EU funded research project ARCOMEM (ICT 270239). 

This program is free software; you can redistribute it and/or modify
it under the terms of version 3 of the GNU General Public License as
published by the Free Software Foundation.

For question or bug reports please contact 
Nina N. Tahmasebi <tahmasebi@L3S.de>

---------------------------------
Background
----------
The ARCOMEM Named Entity Evolution Recognizer, here on called NEER, 
uses time periods of high likelyhood of named entity evolution as starting point 
for detecting evolution for a given query term q. These time periods are 
called <b>change periods<\b> and can be approximated using burst detection
algorithms, for example [1]. All documents from these change periods that mention
the query term q = q1 q2 q3 or any part of the term qi for i = 1, 2, 3 are extracted.
We use a sliding window co-occurrence method with a window size of 20 
(10 words on each side of q) to create a co-occurrence graph (also called context). 
By analyzing only these time periods using a sliding window co-occurrence method 
we capture evolving terms in the same context. We thus avoid comparing terms from 
widely different periods in time. 

We define <b>co-references<\b> are expressions that refer to the same entity. 
In the sentence The president said he had discussed the issue the words 
<b>the president<\b> and <b>he<\b> refer to the same person. For NEER,
we consider temporal co-references to be different lexical representations that have
been used to reference the same concept or entity at the different periods in time.
We have two variations of temporal co-references, direct and indirect. Direct temporal coreferences
are temporal co-references that are variations of each other with some lexical
overlap. Indirect temporal co-references are temporal co-references that lack lexical overlap
on the token level. <b>Hillary Clinton<\b> and <b>Hillary Rodham<\b> are examples of direct 
temporal coreferences while <b>Pope Benedict XVI<\b> and <b>Joseph Ratzinger<\b> are examples 
of indirect temporal co-references. All introduced terms will be used with and without 
temporal interchangeably. A temporal co-reference class contains all direct temporal 
co-references for a given named entity, denoted as coref_r{w1,w2, . . .}. Each temporal 
co-reference class is represented by a class representative r which is also a member of the class. 
For example, Joseph Ratzinger is the representative of the co-reference class containing 
the terms {Joseph Ratzinger, Cardinal Ratzinger, Cardinal Joseph Ratzinger, . . . }.

We use a set of rules in order to find all direct co-references given a <b>dictionary<\b> 
of extracted terms (nouns of length 3) and named entities (extracted using a NER, e.g., [2]).
In the initial iteration the first rule works on the dictionary terms and populates an 
index with co-reference  representatives. In the second and all subsequent iterations the 
first rule makes use of the terms in the index. This index is passed through all the rules.
The rules are iterated until there are no more terms in the index that can be merged.

1. <b>Prefix/suffix rule<b>: This rule creates co-reference classes by merging dictionary terms that
differ only by a prefix or suffix. For example, the co-reference classes of Pope Benedict and
Benedict as well as Pope and Pope Benedict are merged. In both cases the co-reference
class has Pope Benedict as representative and these co-reference classes are therefore
merged and result in coref_Pope Benedict{Pope, Pope Benedict, Benedict}.


2. <b>Sub-term rule<b>: This rule merges classes that are represented by terms that can be considered
sub-terms. For a term to be a sub-term of another we require the longer term to contain
all terms from the shorter term in the correct order. For example, the co-reference classes
represented by Cardinal Joseph Ratzinger and Cardinal Ratzinger are merged.

3. <b>Prolong rule<b>: The third rule is used to create longer terms than might be found in the
dictionary. It merges two representatives from the index into one longer term if the terms
have an overlapping part and there exists a co-occurrence between the remaining terms.
E.g., Pope John Paul and John Paul II are merged if there is a co-occurrence (Pope John
Paul , II) or (Pope , John Paul II); the representative of the merged co-reference class is
Pope John Paul II.  The third rule also merges terms that differ due to plural of the prefixes
assuming that the prefix is not considered a stopword. E.g., Senator Barack Obama and
Senators Barack Obama are merged but Mr Obama and Mrs Obama are not.


<b>Final merging<\b> When the terms in the index cannot be merged further, a final round of
merging takes place. In this round we apply a soft sub-term rule where we drop the requirement
that the terms should be in the same order but require them to be similar in frequency. This
way terms like Illinois Democrat and Democrat of Illinois are merged.


<b>Consolidation<\b> When all terms are merged we create a mapping from each term to the coreference
class representative that has the highest frequency. Using this map we consolidate all
terms in the context of each class.


<b>Indirect Co-references<\b> Indirect co-references are found implicitly by means of the direct
co-references. After consolidation, all terms in the context (co-occurrence graph) of a 
co-reference class are considered candidate indirect co-references. These are a mix between 
true indirect co-references, highly related co-occurrence phrases as well as noise. 
The quality of the indirect co-references is dependent on the named entity extraction, 
co-occurrence graph creation and filtering of the co-occurrence graph. The choice of including 
single token terms in addition to multi-token terms has a high influence on the quality of 
the resulting co-occurrences.


More information and examples for NEER can be found in [3].




The dictionary should be in the below format, with a tab 
separated term and frequency:
------------------------------------------------------------
E.g., 
Pope	993
Benedict	768
John Paul II	708
John	605
Paul	576
Paul II	418
Pope John Paul II	385
XVI	311
Benedict XVI	298
Pope Benedict XVI	281
------------------------------------------------------------

The co-occurrence graph (i.e., the context) should have the following
format: with two terms, the total frequency with which they co-occur and
the normalized frequency with which they co-occur, all tab separated.

------------------------------------------------------------
E.g., 
Pope_John_Paul_II	Vatican	27	1.000000
Benedict	John_Paul	23	0.851852
Congregation	Doctrine	20	0.740741
Congregation	Faith	19	0.703704
Doctrine	Faith	19	0.703704
Pope_Benedict_XVI	Vatican	17	0.629630
Cardinal_Joseph_Ratzinger	Pope_Benedict_XVI	16	0.592593
Pope	Pope_John_Paul_II	15	0.555556
Pope_John_Paul_II	Rome	14	0.518519
Cardinal	Pope_John_Paul_II	12	0.444444
Christ	Church	12	0.444444
MAGIC_TREE_HOUSE	Mary_Pope_Osborne	12	0.444444
Pope	Pope_Benedict_XVI	12	0.444444
------------------------------------------------------------

The output of NEER looks has the following format:
Query term : relative score for the term 
[direct co-references] [indirect co-references]

------------------------------------------------------------
Cardinal_Ratzinger : 17587.0 
[Joseph_Ratzinger, Joseph_Cardinal_Ratzinger, Cardinal_Joseph_Ratzinger, Cardinal_Josef_Ratzinger] 
[Cardinal_Eduardo_Martnez_Somalo, Cardinal_Christoph_Schnborn, Cardinal_Francis_George, 
Cardinal_Mahony, Cardinal_Hayes_High_School, Cardinal_O, Pope_John_Paul_II, Pope_Benedict_XVI, 
Vatican, Congregation, Doctrine, Catholic, Catholics, Faith]
------------------------------------------------------------




Usage:
------
java -jar ARCOMEM-NEER.v2.jar <dictionary_location> <graph_location> <query_term> 
							<stopwordsfile> <burst1> <burst2> ...
where:
	<dictionary_location>	directory for dictionary files
	<graph_location> 		directory for graph files
	<query_term> 			query term where each space is replaces with _
	<stopwordsfile>			directory and file with all stopwords
	<bursts>				as many bursts as needed, e.g., 1924 or Dec2008
	
	It is required that the dictionary- and graphnames contain the query term and the bursts
	and that there exist one dictionary and one graph per burst.
	E.g., dictionary: 		/home/Arcomem/dictionaries/dict_Pope-Benedict-XVI.2005_Lingua_NP_v2.txt 
		  graph:			/home/Arcomem/graphs/SlidingWindowGraph_Pope-Benedict-XVI.2005_Lingua_NP_v2.txt
		  query:			Pope_Benedict_XVI
		  stopwordsfile:	/home/Acromem/stopwords.txt
		  burst:			2005
	
Example usage:
java -jar ARCOMEM-NEER.v2.jar ~\ARCOMEM-NEER\data\dictionaries/ ~\ARCOMEM-NEER\data\graphs/ Cardinal_Ratzinger ~\ARCOMEM-NEER\data\stopwords.txt 2005 2001
Expected output:
---------------
Cardinal_Ratzinger : 17587.0 [Joseph_Ratzinger, Joseph_Cardinal_Ratzinger, Cardinal_Joseph_Ratzinger, Cardinal_Josef_Ratzinger] 
[Cardinal_Eduardo_Martnez_Somalo, Cardinal_Christoph_Schnborn, Cardinal_Francis_George, Cardinal_Mahony, Cardinal_Hayes_High_School,
 Cardinal_O, Pope_John_Paul_II, Pope_Benedict_XVI, Vatican, Congregation, Doctrine, Catholic, Catholics, Faith]

 
--------------------------------------------
	[1] Kleinberg, J. (2003). Bursty and hierarchical structure in streams. 
		Data Mining and Knowledge Discovery, 7(4):373397.
	[2] Finkel, J. R., Grenager, T., and Manning, C. (2005). Incorporating 
	non-local information into information extraction systems by Gibbs sampling. 
	In ACL, pages 363370.	
	[3] Nina Tahmasebi, Gerhard Gossen, Nattiya Kanhabua, Helge Holzmann, 
		Thomas Risse: NEER: An Unsupervised Method for Named Entity Evolution 
		Recognition. COLING 2012: 2553-2568