summaryrefslogtreecommitdiff
path: root/examples/server/tests/features/parallel.feature
diff options
context:
space:
mode:
Diffstat (limited to 'examples/server/tests/features/parallel.feature')
-rw-r--r--examples/server/tests/features/parallel.feature5
1 files changed, 3 insertions, 2 deletions
diff --git a/examples/server/tests/features/parallel.feature b/examples/server/tests/features/parallel.feature
index 066698c8..a66fed62 100644
--- a/examples/server/tests/features/parallel.feature
+++ b/examples/server/tests/features/parallel.feature
@@ -6,8 +6,8 @@ Feature: Parallel
Given a server listening on localhost:8080
And a model file tinyllamas/stories260K.gguf from HF repo ggml-org/models
And 42 as server seed
- And 512 as batch size
- And 64 KV cache size
+ And 128 as batch size
+ And 256 KV cache size
And 2 slots
And continuous batching
Then the server is starting
@@ -76,6 +76,7 @@ Feature: Parallel
| disabled | 128 |
| enabled | 64 |
+
Scenario: Multi users with total number of tokens to predict exceeds the KV Cache size #3969
Given a prompt:
"""