Abstract:
Because some problems existed in traditional token-based algorithm for homology detection in structured information location, module identification, module extraction and high precision homology measure for code variants, a structured recognition homology detection technology was proposed based on an improved edit distance algorithm and improved longest common sequence (LCS) algorithm. In the edit distance calculation, the exchange operator was introduced to improve the measurement accuracy of internal homology modules. In the LCS algorithm, a minimum size monitoring mechanism and line maximum dynamic correlation measure were introduced for similar modules, which offered the ability of code structure boundary division, module line association and structured information extraction. Experiments show that the structure information based algorithm is effective and stable for code homology detection, and the results of random sampling detection show its better performances in precision, recall rate and
F values. Experiments show that the algorithm utilizing structure information for code homology detection is effective and stable, and the results of random sampling detection have better performances in precision, recall rate and
F values.