Compare commits

..
3 Commits
Author SHA1 Message Date
hz4th_coder 4f02a6a0ea 实现浏览器方式互联网搜索功能
- 使用 agent-browser 浏览器自动化进行 Bing 搜索
- 解析 accessibility tree 提取搜索结果
- 自动获取每个结果的 URL
- 支持设置最大结果数量
2026-07-13 11:55:33 +08:00
hz4th_coder 187d5037b8 添加Git推送说明文档 2026-07-12 01:24:46 +08:00
hz4th_coder 47729a52da 添加部署文档和说明 2026-07-12 01:24:19 +08:00
3 changed files with 523 additions and 5 deletions
+302
View File
@@ -0,0 +1,302 @@
# Param Auto Manager 部署文档
## 📦 系统概览
**参数数据自动化管理系统** - 自动管理参数数据网站的系统,支持从内容库和互联网搜索数据,自动处理产品并提交到后台管理待审核区。
---
## 🚀 快速部署
### 1. 系统要求
- Python 3.10+
- Flask
- SQLite3
- Git
### 2. 安装依赖
```bash
cd works/param-auto-manager
pip install -r requirements.txt
```
### 3. 启动服务
```bash
# 方式1: 使用启动脚本
./start.sh
# 方式2: 直接运行
export PYTHONPATH="$HOME/.local/lib/python3.12/site-packages:$PYTHONPATH"
python3 app.py
```
### 4. 访问系统
- **前端界面**: http://localhost:16043 或 http://192.168.0.101:16043
- **API接口**: http://localhost:16043/api
---
## 🌐 前端界面功能
### 主要功能模块
1. **统计面板**
- 内容库文章数量
- 待处理产品数量
- 处理中产品数量
- 最近成功/失败数量
2. **待处理产品列表**
- 查看所有待处理产品
- 添加单个/批量产品
- 手动触发处理
- 移除产品
- 批量处理按钮
3. **内容库文章**
- 查看所有文章
- 搜索文章(关键词)
- 添加新文章
- 查看文章详情
- 删除文章
4. **处理历史**
- 查看处理历史记录
- 状态显示(已提交/失败/错误)
- 查看处理详情
5. **系统配置**
- 自动处理开关
- 处理间隔设置
- 批量处理数量
---
## 🔌 API接口
### 文章内容库 API
```bash
# 获取文章列表
GET /api/articles
# 搜索文章
GET /api/articles/search?q=关键词
# 创建文章
POST /api/articles
{
"product_names": ["产品1", "产品2"],
"category": "分类",
"keywords": ["关键词"],
"summary": "摘要",
"content": "内容",
"source": "来源"
}
# 删除文章
DELETE /api/articles/{id}
```
### 产品处理 API
```bash
# 待处理产品列表
GET /api/products/pending
# 添加产品
POST /api/products/pending
{
"product_name": "产品名称",
"category": "分类",
"priority": 10
}
# 处理单个产品
POST /api/products/process
{
"product_name": "产品名称"
}
# 批量处理
POST /api/products/process/batch
{
"limit": 5
}
# 处理历史
GET /api/products/history
```
### 系统管理 API
```bash
# 系统统计
GET /api/system/stats
# 系统配置
GET /api/system/config
PUT /api/system/config
# 健康检查
GET /api/system/health
```
---
## ⚙️ Git仓库推送
### ⚠️ 注意事项
当前Git服务器不允许用户直接推送创建仓库(Push to create),需要管理员预先创建仓库。
### 推送步骤
**方式1: 管理员预先创建仓库**
1. 在Git服务器创建仓库:
- 地址: `http://121.40.164.32:12007/hz4th_coder/param-auto-manager.git`
- 账号: `hz4th_coder`
2. 推送代码:
```bash
cd works/param-auto-manager
git remote add origin http://hz4th_coder:262e7dbfce09c8cc21fbacff2b450cc3f1c3e265@121.40.164.32:12007/hz4th_coder/param-auto-manager.git
git push -u origin master
git push origin v1.1.0
```
**方式2: 请求管理员推送权限**
联系Git服务器管理员,开通用户的"Push to create"权限。
---
## 📊 使用示例
### 1. 通过前端界面操作
1. **添加待处理产品**
- 点击"添加产品"按钮
- 输入产品名称(如: GPT-4o
- 设置分类和优先级
- 点击"添加"
2. **添加文章到内容库**
- 点击"添加文章"按钮
- 填写产品名称、摘要、内容等
- 点击"添加"
3. **手动触发处理**
- 在待处理产品列表找到产品
- 点击"处理"按钮
- 查看处理结果
### 2. 通过API操作
```bash
# 添加产品
curl -X POST http://localhost:16043/api/products/pending \
-H "Content-Type: application/json" \
-d '{"product_name":"GPT-4o","category":"AI模型","priority":10}'
# 添加文章
curl -X POST http://localhost:16043/api/articles \
-H "Content-Type: application/json" \
-d '{
"product_names": ["GPT-4o"],
"category": "AI模型",
"keywords": ["大模型","多模态"],
"summary": "OpenAI多模态大模型",
"content": "详细内容...",
"source": "官方网站"
}'
# 触发处理
curl -X POST http://localhost:16043/api/products/process \
-H "Content-Type: application/json" \
-d '{"product_name":"GPT-4o"}'
```
---
## 🔧 系统配置
配置文件: `config.py`
```python
PORT = 16043 # 服务端口
PARAMHUB_BASE_URL = 'http://localhost:16041' # ParamHub API地址
PROCESS_INTERVAL = 300 # 自动处理间隔(秒)
BATCH_SIZE = 5 # 批量处理数量
```
---
## 📁 项目结构
```
param-auto-manager/
├── app.py # 主应用
├── config.py # 配置文件
├── models/
│ └── database.py # 数据库模型
├── routes/
│ ├── articles.py # 文章API
│ ├── products.py # 产品API
│ └── system.py # 系统API
├── services/
│ ├── search_service.py # 搜索服务
│ ├── process_service.py # 处理服务
│ └── paramhub_client.py # ParamHub客户端
├── templates/
│ └── index.html # 前端页面
├── static/
│ ├── css/style.css # 样式文件
│ └── js/app.js # JavaScript文件
├── utils/
│ └── scheduler.py # 定时任务
├── data/ # 数据库文件
└── logs/ # 日志文件
```
---
## ✅ 当前部署状态
- **端口**: 16043 ✅
- **状态**: 运行中
- **PID**: 2710498
- **前端**: http://192.168.0.101:16043 ✅
- **API**: http://192.168.0.101:16043/api ✅
- **数据库**: SQLite (param_auto.db) ✅
- **定时任务**: 每5分钟自动处理 ✅
---
## 📝 版本历史
- **v1.1.0** (2026-07-12): 添加前端界面和操作页面
- **v1.0.0** (2026-07-12): 初始版本,核心功能实现
---
## 🆘 常见问题
### Q: Git推送失败怎么办?
A: 需要在Git服务器上预先创建仓库,或联系管理员开通推送权限。
### Q: 如何查看日志?
A: 日志文件位于 `logs/app.log`
### Q: 如何停止服务?
A: 运行 `./stop.sh` 或手动杀掉进程
### Q: 前端页面打不开?
A: 检查服务是否启动,查看日志是否有错误
---
## 📞 支持
如有问题,请查看日志文件或联系系统管理员。
+88
View File
@@ -0,0 +1,88 @@
# Git推送说明
## 📌 当前状态
代码已准备推送,但Git服务器返回403错误:
```
remote: Push to create is not enabled for users.
```
这表示Git服务器不允许用户直接推送创建仓库。
---
## ✅ 解决方案
### 方案1: 管理员预先创建仓库
请在Git服务器上手动创建以下仓库:
**仓库信息:**
- **地址**: http://121.40.164.32:12007/hz4th_coder/param-auto-manager.git
- **账号**: hz4th_coder
- **组织**: hz4th_coder
**创建后推送代码:**
```bash
cd /home/openclaw/.openclaw/workspace-hz4th_coder/works/param-auto-manager
# 添加远程仓库(如果还没有)
git remote add origin http://hz4th_coder:262e7dbfce09c8cc21fbacff2b450cc3f1c3e265@121.40.164.32:12007/hz4th_coder/param-auto-manager.git
# 推送代码和标签
git push -u origin master
git push origin v1.0.0
git push origin v1.1.0
git push origin v1.2.0
```
---
### 方案2: 开通推送权限
联系Git服务器管理员,为 `hz4th_coder` 用户开通"Push to create"权限。
---
## 📊 当前Git状态
```bash
cd works/param-auto-manager
git log --oneline -5
```
输出:
```
47729a5 添加部署文档和说明
8ad0246 添加前端界面和操作页面
e3bc883 添加.gitignore文件,排除缓存和临时文件
b7b0925 初始化参数数据自动化管理系统
```
标签:
```
v1.0.0 - 初始化版本
v1.1.0 - 添加前端界面
v1.2.0 - 添加部署文档
```
---
## 🔐 Git认证信息
- **账号**: hz4th_coder
- **邮箱**: hz4th_coder@tphai.com
- **Token**: 262e7dbfce09c8cc21fbacff2b450cc3f1c3e265
---
## 📝 待推送文件统计
- 总文件数: 26个源文件
- 总代码行数: 约5000行
- 版本标签: 3个
---
**建议**: 请在Git服务器创建仓库后,执行推送命令即可完成部署。
+133 -5
View File
@@ -4,6 +4,10 @@
import requests
from bs4 import BeautifulSoup
import json
import subprocess
import os
import re
import urllib.parse
from datetime import datetime
from config import Config
from models.database import db
@@ -13,19 +17,143 @@ class SearchService:
self.timeout = Config.SEARCH_TIMEOUT
self.max_results = Config.SEARCH_MAX_RESULTS
def _run_browser(self, *args, timeout=30000):
"""运行 agent-browser 命令"""
env = os.environ.copy()
env['XDG_RUNTIME_DIR'] = '/tmp/agent-browser-runtime'
os.makedirs(env['XDG_RUNTIME_DIR'], exist_ok=True)
cmd = ['agent-browser'] + list(args)
result = subprocess.run(
cmd,
capture_output=True,
text=True,
env=env,
timeout=timeout // 1000 + 5
)
return result.stdout, result.stderr, result.returncode
def search_internet(self, keyword, max_results=None):
"""
从互联网搜索(使用搜索API或爬虫
这里暂时使用简单的搜索模拟
从互联网搜索(使用 agent-browser 浏览器自动化
"""
max_results = max_results or self.max_results
# TODO: 接入真实的搜索API(如Google Custom Search、Bing等)
# 这里先返回空列表,等待后续接入真实API
results = []
try:
# 1. 打开 Bing 搜索
encoded_keyword = urllib.parse.quote(keyword)
search_url = f"https://www.bing.com/search?q={encoded_keyword}"
stdout, stderr, code = self._run_browser('open', search_url, '--timeout', '20000')
if code != 0:
print(f"打开搜索页面失败: {stderr}")
return results
# 等待页面加载
stdout, stderr, code = self._run_browser('wait', '5000')
# 2. 获取搜索结果页面结构 (JSON 格式)
stdout, stderr, code = self._run_browser('snapshot', '--json', '--timeout', '30000')
if code != 0:
print(f"获取页面结构失败: {stderr}")
return results
# 3. 解析 JSON 提取搜索结果
try:
data = json.loads(stdout)
except json.JSONDecodeError:
print(f"解析 JSON 失败: {stdout[:500]}")
return results
# 4. 从 accessibility tree 中提取搜索结果
# Bing 搜索结果在 main[aria-label="搜索结果"] 区域内
results = self._parse_bing_results(data, max_results)
# 5. 关闭浏览器
self._run_browser('close')
except subprocess.TimeoutExpired:
print(f"搜索超时: {keyword}")
except Exception as e:
print(f"搜索出错: {str(e)}")
# 尝试关闭浏览器
try:
self._run_browser('close')
except:
pass
return results
def _parse_bing_results(self, snapshot_data, max_results=10):
"""
从 Bing 搜索结果的 snapshot 中解析出标题和链接
snapshot_data 是 agent-browser snapshot --json 的输出
结构: {success, data: {snapshot: "文本格式的 accessibility tree"}, error}
"""
results = []
# 获取 snapshot 文本
snapshot = snapshot_data.get('data', {}).get('snapshot', '')
if not snapshot:
return results
# 解析 accessibility tree 文本
in_results = False
refs = [] # 存储 (title, ref) 元组
lines = snapshot.split('\n')
for i, line in enumerate(lines):
line = line.strip()
# 进入搜索结果区域
if 'main "搜索结果"' in line:
in_results = True
continue
# 离开搜索结果区域
if in_results and line.startswith('- ') and 'main' in line and '搜索结果' not in line:
break
if not in_results:
continue
# 匹配标题链接:link "标题文字" [ref=eXX]
# 需要过滤域名链接(如 "zhihu.com")和短链接
if 'link "' in line and '[ref=' in line:
match = re.search(r'link "([^"]+)" \[ref=(e\d+)\]', line)
if match:
title = match.group(1)
ref = match.group(2)
# 过滤短标题(域名链接如 "zhihu.com"
if len(title) > 20 and '.' not in title[:10]: # 不是域名格式
refs.append((title, ref))
# 获取每个结果的 URL
for title, ref in refs[:max_results]:
url = self._get_link_url(ref)
if url and 'bing.com/search' not in url: # 过滤搜索结果页本身的链接
results.append({
'title': title,
'url': url,
'snippet': '',
'source': 'bing'
})
return results
def _get_link_url(self, ref):
"""通过 agent-browser 获取链接的 URL"""
try:
stdout, stderr, code = self._run_browser('get', 'attr', f'@{ref}', 'href', '--json', '--timeout', '5000')
if code == 0 and stdout:
data = json.loads(stdout)
return data.get('data', {}).get('value', '')
except Exception as e:
print(f"获取 URL 失败 (ref={ref}): {e}")
return None
def fetch_url_content(self, url):
"""抓取网页内容"""
try: