Compare commits

...
3 Commits
Author SHA1 Message Date
hz4th_coder 4f02a6a0ea 实现浏览器方式互联网搜索功能
- 使用 agent-browser 浏览器自动化进行 Bing 搜索
- 解析 accessibility tree 提取搜索结果
- 自动获取每个结果的 URL
- 支持设置最大结果数量
2026-07-13 11:55:33 +08:00
hz4th_coder 187d5037b8 添加Git推送说明文档 2026-07-12 01:24:46 +08:00
hz4th_coder 47729a52da 添加部署文档和说明 2026-07-12 01:24:19 +08:00
3 changed files with 523 additions and 5 deletions
+302
View File
@@ -0,0 +1,302 @@
# Param Auto Manager 部署文档
## 📦 系统概览
**参数数据自动化管理系统** - 自动管理参数数据网站的系统,支持从内容库和互联网搜索数据,自动处理产品并提交到后台管理待审核区。
---
## 🚀 快速部署
### 1. 系统要求
- Python 3.10+
- Flask
- SQLite3
- Git
### 2. 安装依赖
```bash
cd works/param-auto-manager
pip install -r requirements.txt
```
### 3. 启动服务
```bash
# 方式1: 使用启动脚本
./start.sh
# 方式2: 直接运行
export PYTHONPATH="$HOME/.local/lib/python3.12/site-packages:$PYTHONPATH"
python3 app.py
```
### 4. 访问系统
- **前端界面**: http://localhost:16043 或 http://192.168.0.101:16043
- **API接口**: http://localhost:16043/api
---
## 🌐 前端界面功能
### 主要功能模块
1. **统计面板**
- 内容库文章数量
- 待处理产品数量
- 处理中产品数量
- 最近成功/失败数量
2. **待处理产品列表**
- 查看所有待处理产品
- 添加单个/批量产品
- 手动触发处理
- 移除产品
- 批量处理按钮
3. **内容库文章**
- 查看所有文章
- 搜索文章(关键词)
- 添加新文章
- 查看文章详情
- 删除文章
4. **处理历史**
- 查看处理历史记录
- 状态显示(已提交/失败/错误)
- 查看处理详情
5. **系统配置**
- 自动处理开关
- 处理间隔设置
- 批量处理数量
---
## 🔌 API接口
### 文章内容库 API
```bash
# 获取文章列表
GET /api/articles
# 搜索文章
GET /api/articles/search?q=关键词
# 创建文章
POST /api/articles
{
"product_names": ["产品1", "产品2"],
"category": "分类",
"keywords": ["关键词"],
"summary": "摘要",
"content": "内容",
"source": "来源"
}
# 删除文章
DELETE /api/articles/{id}
```
### 产品处理 API
```bash
# 待处理产品列表
GET /api/products/pending
# 添加产品
POST /api/products/pending
{
"product_name": "产品名称",
"category": "分类",
"priority": 10
}
# 处理单个产品
POST /api/products/process
{
"product_name": "产品名称"
}
# 批量处理
POST /api/products/process/batch
{
"limit": 5
}
# 处理历史
GET /api/products/history
```
### 系统管理 API
```bash
# 系统统计
GET /api/system/stats
# 系统配置
GET /api/system/config
PUT /api/system/config
# 健康检查
GET /api/system/health
```
---
## ⚙️ Git仓库推送
### ⚠️ 注意事项
当前Git服务器不允许用户直接推送创建仓库(Push to create),需要管理员预先创建仓库。
### 推送步骤
**方式1: 管理员预先创建仓库**
1. 在Git服务器创建仓库:
- 地址: `http://121.40.164.32:12007/hz4th_coder/param-auto-manager.git`
- 账号: `hz4th_coder`
2. 推送代码:
```bash
cd works/param-auto-manager
git remote add origin http://hz4th_coder:262e7dbfce09c8cc21fbacff2b450cc3f1c3e265@121.40.164.32:12007/hz4th_coder/param-auto-manager.git
git push -u origin master
git push origin v1.1.0
```
**方式2: 请求管理员推送权限**
联系Git服务器管理员,开通用户的"Push to create"权限。
---
## 📊 使用示例
### 1. 通过前端界面操作
1. **添加待处理产品**
- 点击"添加产品"按钮
- 输入产品名称(如: GPT-4o
- 设置分类和优先级
- 点击"添加"
2. **添加文章到内容库**
- 点击"添加文章"按钮
- 填写产品名称、摘要、内容等
- 点击"添加"
3. **手动触发处理**
- 在待处理产品列表找到产品
- 点击"处理"按钮
- 查看处理结果
### 2. 通过API操作
```bash
# 添加产品
curl -X POST http://localhost:16043/api/products/pending \
-H "Content-Type: application/json" \
-d '{"product_name":"GPT-4o","category":"AI模型","priority":10}'
# 添加文章
curl -X POST http://localhost:16043/api/articles \
-H "Content-Type: application/json" \
-d '{
"product_names": ["GPT-4o"],
"category": "AI模型",
"keywords": ["大模型","多模态"],
"summary": "OpenAI多模态大模型",
"content": "详细内容...",
"source": "官方网站"
}'
# 触发处理
curl -X POST http://localhost:16043/api/products/process \
-H "Content-Type: application/json" \
-d '{"product_name":"GPT-4o"}'
```
---
## 🔧 系统配置
配置文件: `config.py`
```python
PORT = 16043 # 服务端口
PARAMHUB_BASE_URL = 'http://localhost:16041' # ParamHub API地址
PROCESS_INTERVAL = 300 # 自动处理间隔(秒)
BATCH_SIZE = 5 # 批量处理数量
```
---
## 📁 项目结构
```
param-auto-manager/
├── app.py # 主应用
├── config.py # 配置文件
├── models/
│ └── database.py # 数据库模型
├── routes/
│ ├── articles.py # 文章API
│ ├── products.py # 产品API
│ └── system.py # 系统API
├── services/
│ ├── search_service.py # 搜索服务
│ ├── process_service.py # 处理服务
│ └── paramhub_client.py # ParamHub客户端
├── templates/
│ └── index.html # 前端页面
├── static/
│ ├── css/style.css # 样式文件
│ └── js/app.js # JavaScript文件
├── utils/
│ └── scheduler.py # 定时任务
├── data/ # 数据库文件
└── logs/ # 日志文件
```
---
## ✅ 当前部署状态
- **端口**: 16043 ✅
- **状态**: 运行中
- **PID**: 2710498
- **前端**: http://192.168.0.101:16043 ✅
- **API**: http://192.168.0.101:16043/api ✅
- **数据库**: SQLite (param_auto.db) ✅
- **定时任务**: 每5分钟自动处理 ✅
---
## 📝 版本历史
- **v1.1.0** (2026-07-12): 添加前端界面和操作页面
- **v1.0.0** (2026-07-12): 初始版本,核心功能实现
---
## 🆘 常见问题
### Q: Git推送失败怎么办?
A: 需要在Git服务器上预先创建仓库,或联系管理员开通推送权限。
### Q: 如何查看日志?
A: 日志文件位于 `logs/app.log`
### Q: 如何停止服务?
A: 运行 `./stop.sh` 或手动杀掉进程
### Q: 前端页面打不开?
A: 检查服务是否启动,查看日志是否有错误
---
## 📞 支持
如有问题,请查看日志文件或联系系统管理员。
+88
View File
@@ -0,0 +1,88 @@
# Git推送说明
## 📌 当前状态
代码已准备推送,但Git服务器返回403错误:
```
remote: Push to create is not enabled for users.
```
这表示Git服务器不允许用户直接推送创建仓库。
---
## ✅ 解决方案
### 方案1: 管理员预先创建仓库
请在Git服务器上手动创建以下仓库:
**仓库信息:**
- **地址**: http://121.40.164.32:12007/hz4th_coder/param-auto-manager.git
- **账号**: hz4th_coder
- **组织**: hz4th_coder
**创建后推送代码:**
```bash
cd /home/openclaw/.openclaw/workspace-hz4th_coder/works/param-auto-manager
# 添加远程仓库(如果还没有)
git remote add origin http://hz4th_coder:262e7dbfce09c8cc21fbacff2b450cc3f1c3e265@121.40.164.32:12007/hz4th_coder/param-auto-manager.git
# 推送代码和标签
git push -u origin master
git push origin v1.0.0
git push origin v1.1.0
git push origin v1.2.0
```
---
### 方案2: 开通推送权限
联系Git服务器管理员,为 `hz4th_coder` 用户开通"Push to create"权限。
---
## 📊 当前Git状态
```bash
cd works/param-auto-manager
git log --oneline -5
```
输出:
```
47729a5 添加部署文档和说明
8ad0246 添加前端界面和操作页面
e3bc883 添加.gitignore文件,排除缓存和临时文件
b7b0925 初始化参数数据自动化管理系统
```
标签:
```
v1.0.0 - 初始化版本
v1.1.0 - 添加前端界面
v1.2.0 - 添加部署文档
```
---
## 🔐 Git认证信息
- **账号**: hz4th_coder
- **邮箱**: hz4th_coder@tphai.com
- **Token**: 262e7dbfce09c8cc21fbacff2b450cc3f1c3e265
---
## 📝 待推送文件统计
- 总文件数: 26个源文件
- 总代码行数: 约5000行
- 版本标签: 3个
---
**建议**: 请在Git服务器创建仓库后,执行推送命令即可完成部署。
+133 -5
View File
@@ -4,6 +4,10 @@
import requests import requests
from bs4 import BeautifulSoup from bs4 import BeautifulSoup
import json import json
import subprocess
import os
import re
import urllib.parse
from datetime import datetime from datetime import datetime
from config import Config from config import Config
from models.database import db from models.database import db
@@ -13,19 +17,143 @@ class SearchService:
self.timeout = Config.SEARCH_TIMEOUT self.timeout = Config.SEARCH_TIMEOUT
self.max_results = Config.SEARCH_MAX_RESULTS self.max_results = Config.SEARCH_MAX_RESULTS
def _run_browser(self, *args, timeout=30000):
"""运行 agent-browser 命令"""
env = os.environ.copy()
env['XDG_RUNTIME_DIR'] = '/tmp/agent-browser-runtime'
os.makedirs(env['XDG_RUNTIME_DIR'], exist_ok=True)
cmd = ['agent-browser'] + list(args)
result = subprocess.run(
cmd,
capture_output=True,
text=True,
env=env,
timeout=timeout // 1000 + 5
)
return result.stdout, result.stderr, result.returncode
def search_internet(self, keyword, max_results=None): def search_internet(self, keyword, max_results=None):
""" """
从互联网搜索(使用搜索API或爬虫 从互联网搜索(使用 agent-browser 浏览器自动化
这里暂时使用简单的搜索模拟
""" """
max_results = max_results or self.max_results max_results = max_results or self.max_results
# TODO: 接入真实的搜索API(如Google Custom Search、Bing等)
# 这里先返回空列表,等待后续接入真实API
results = [] results = []
try:
# 1. 打开 Bing 搜索
encoded_keyword = urllib.parse.quote(keyword)
search_url = f"https://www.bing.com/search?q={encoded_keyword}"
stdout, stderr, code = self._run_browser('open', search_url, '--timeout', '20000')
if code != 0:
print(f"打开搜索页面失败: {stderr}")
return results
# 等待页面加载
stdout, stderr, code = self._run_browser('wait', '5000')
# 2. 获取搜索结果页面结构 (JSON 格式)
stdout, stderr, code = self._run_browser('snapshot', '--json', '--timeout', '30000')
if code != 0:
print(f"获取页面结构失败: {stderr}")
return results
# 3. 解析 JSON 提取搜索结果
try:
data = json.loads(stdout)
except json.JSONDecodeError:
print(f"解析 JSON 失败: {stdout[:500]}")
return results
# 4. 从 accessibility tree 中提取搜索结果
# Bing 搜索结果在 main[aria-label="搜索结果"] 区域内
results = self._parse_bing_results(data, max_results)
# 5. 关闭浏览器
self._run_browser('close')
except subprocess.TimeoutExpired:
print(f"搜索超时: {keyword}")
except Exception as e:
print(f"搜索出错: {str(e)}")
# 尝试关闭浏览器
try:
self._run_browser('close')
except:
pass
return results return results
def _parse_bing_results(self, snapshot_data, max_results=10):
"""
从 Bing 搜索结果的 snapshot 中解析出标题和链接
snapshot_data 是 agent-browser snapshot --json 的输出
结构: {success, data: {snapshot: "文本格式的 accessibility tree"}, error}
"""
results = []
# 获取 snapshot 文本
snapshot = snapshot_data.get('data', {}).get('snapshot', '')
if not snapshot:
return results
# 解析 accessibility tree 文本
in_results = False
refs = [] # 存储 (title, ref) 元组
lines = snapshot.split('\n')
for i, line in enumerate(lines):
line = line.strip()
# 进入搜索结果区域
if 'main "搜索结果"' in line:
in_results = True
continue
# 离开搜索结果区域
if in_results and line.startswith('- ') and 'main' in line and '搜索结果' not in line:
break
if not in_results:
continue
# 匹配标题链接:link "标题文字" [ref=eXX]
# 需要过滤域名链接(如 "zhihu.com")和短链接
if 'link "' in line and '[ref=' in line:
match = re.search(r'link "([^"]+)" \[ref=(e\d+)\]', line)
if match:
title = match.group(1)
ref = match.group(2)
# 过滤短标题(域名链接如 "zhihu.com"
if len(title) > 20 and '.' not in title[:10]: # 不是域名格式
refs.append((title, ref))
# 获取每个结果的 URL
for title, ref in refs[:max_results]:
url = self._get_link_url(ref)
if url and 'bing.com/search' not in url: # 过滤搜索结果页本身的链接
results.append({
'title': title,
'url': url,
'snippet': '',
'source': 'bing'
})
return results
def _get_link_url(self, ref):
"""通过 agent-browser 获取链接的 URL"""
try:
stdout, stderr, code = self._run_browser('get', 'attr', f'@{ref}', 'href', '--json', '--timeout', '5000')
if code == 0 and stdout:
data = json.loads(stdout)
return data.get('data', {}).get('value', '')
except Exception as e:
print(f"获取 URL 失败 (ref={ref}): {e}")
return None
def fetch_url_content(self, url): def fetch_url_content(self, url):
"""抓取网页内容""" """抓取网页内容"""
try: try: