feat(ai): thinking mode switch, vision image input, and API host normalization fixes - #1596
Open
lb1038678031 wants to merge 4 commits into
Open
feat(ai): thinking mode switch, vision image input, and API host normalization fixes#1596lb1038678031 wants to merge 4 commits into
lb1038678031 wants to merge 4 commits into
Conversation
- Unify host normalization between model listing and chat client so both hosts with or without the /v1 suffix resolve identically (fixes upstream NOT_FOUND when a full /v1 base URL is configured). - Trim trailing slashes before building the models URL. - Summarize upstream error bodies instead of echoing the full response. - Prefer the chosen provider's stored key over host-based key lookup. - Return nil client when AI is disabled instead of an invalid client. - Fix default provider list: use Gemini's OpenAI-compatible endpoint and drop Anthropic (not supported by the OpenAI-compatible backend).
Thinking mode (per provider): - New thinking_mode field on the AI provider config. - When enabled, a custom http.RoundTripper merges enable_thinking=true into chat completion request bodies before they are sent. Vision input (per provider): - New vision_enabled field; surfaced to clients via site info ai_vision_enabled so the front end can show the attach button. - Chat requests accept images on the first user message: base64 data URLs or HTTPS links, max 4 per message and 4MB decoded each. - Images are converted into MultiContent parts for vision models. - Conversation history keeps an '[图片]' placeholder instead of image data. Also stops logging full request bodies that contain image data.
- Admin AI settings page gains deep-thinking and image-input switches per provider, saved with the rest of the provider config. - Site settings store tracks ai_vision_enabled. - AiAssistant sender shows an attach-image action when vision is on, previews thumbnails, validates count/type/size client-side and sends base64 images with the first message. New-conversation first hop now carries images through as well. - zh_CN/en_US translations for all new strings.
The embedded base64 constant in the vision test was mistaken for a leaked credential. Build the placeholder image from the canonical PNG signature bytes at runtime instead, so no base64 blob appears in the source.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resubmit of #1594, addressing @LinkinStars's feedback there. Implements #1595 (feature clarification).
What
Three improvements to the AI feature, shipped as independent commits:
fix(ai): normalize API host and harden AI model listing
api_host + "/v1/models"while the chatclient auto-appends
/v1. Configuring a base URL that already contains/v1(common with OpenAI-compatible gateways such as SenseNova, SiliconFlow,
OneAPI relays...) produced
/v1/v1/models, and the gateway's gRPC-styleNOT_FOUNDerror was passed through to the UI.schema.NormalizeAPIHost()trims slashes/whitespace,appends
/v1when missing, and keeps/v1beta/*endpoints intact.error.code/message) instead of beingechoed verbatim; image data is never logged.
matching.
Anthropic removed because the backend only speaks the OpenAI protocol.
feat(ai): per-provider thinking mode
thinking_modeswitch on each provider in Admin → AI settings.http.RoundTrippermerges"enable_thinking": trueinto chat completion request bodies — the convention used by
reasoning-capable OpenAI-compatible gateways (DeepSeek V4, Qwen/DashScope,
SenseNova, vLLM/SGLang, ...). The existing
reasoning_contentstreamingpath renders the thought process in the bubble UI unchanged.
feat(ai): vision image input behind an admin switch
vision_enabledswitch per provider; exposed to clients via site infoas
ai_vision_enabled.first message of a turn; images travel as base64 data URLs or HTTPS links
and are converted into MultiContent parts.
(no DB schema change).
Security note (from #1594)
The base64 constant flagged during review was the encoding of a standard 1×1
pixel PNG placeholder used in a unit test, not a credential. To avoid any
misreading,
158516bconstructs the placeholder at runtime from the canonicalPNG signature bytes; the source no longer contains any base64 blob.
Testing
go test ./internal/schema/ ./internal/controller/ ./internal/service/siteinfo/ ./internal/migrations/— new table tests for host normalization, transport injection, and image
validation (count/type/size).
succeeds with both root and
/v1base URLs;reasoning_contentappears withthinking enabled; image questions answered by a vision-capable model.