solr系统query检索词特殊字符的处理

来源:互联网 时间:1970-01-01


solr是基于 lucence开发的应用,如果query中带有非法字符串,结果很可能是检索出所有内容或者直接报错,所以你对用户的输入必须要先做处理。输入星号,能够检索出所有内容;输入加号,则会报错。

官方的处理办法(java,因为solr是java开发的):

https://svn.apache.org/repos/asf/lucene/dev/trunk/solr/solrj/src/java/org/apache/solr/client/solrj/util/ClientUtils.javapublic static String escapeQueryChars(String s) { StringBuilder sb = new StringBuilder(); for (int i = 0; i < s.length(); i++) { char c = s.charAt(i); // These characters are part of the query syntax and must be escaped if (c == '//' || c == '+' || c == '-' || c == '!' || c == '(' || c == ')' || c == ':' || c == '^' || c == '[' || c == ']' || c == '/"' || c == '{' || c == '}' || c == '~' || c == '*' || c == '?' || c == '|' || c == '&' || c == ';' || c == '/' || Character.isWhitespace(c)) { sb.append('//'); } sb.append(c); } return sb.toString(); }

翻译的php版本(利用preg_replace函数进行正则替换):

static public function escape($value){ //list taken from http://lucene.apache.org/java/docs/queryparsersyntax.html#Escaping%20Special%20Characters $pattern = '/(/+|-|&|/||!|/(|/)|/{|}|/[|]|/^|"|~|/*|/?|:|;|~|//)/'; $replace = '///$1'; return preg_replace($pattern, $replace, $value);}

翻译后的python版本:

import redef escape_solr(word): return re.sub('(///|/+|-|&|/|/||!|/(|/)|/{|}|/[|]|/^|"|~|/*|/?|:|;|/|/~)','///1', word )






相关阅读:
Top