Optimal weight assignment for a Chinese signature file

Tyne Liang, Suh-Yin Lee, Wei-Pang Yang
1996 Information Processing & Management  
The performance of a character-based Chinese text retrieval scheme (the combined scheme) is investigated. In the scheme both the monogram keys (singleton characters) and bigram keys (consecutive character pairs) are encoded into document signatures such that half of the bits in every signature are set. For disyllabic queries, an analytical expression of the false hit rate that accounts for both random false hits and adjacency false hits is proposed. Then optimal monogram and bigram weight
more » ... ments together with the corresponding minimal false hit rate are derived in terms of signature length, storage overhead of the combined scheme, and the occurrence frequency and the association value of a disyllabic query. The theoretical predictions of the optimal weight assignments and the minimal false hit rate are tested and verified in experiments using a real Chinese cot'pus for disyllabic queries of different association values. Satisfactory agreement between the experimental results and theoretical predictions is found.
doi:10.1016/s0306-4573(96)85008-4 fatcat:a4uvlbv3hnbnpha55eabylivsq