방법론: AI나 사람에게 하나의 보편적 accuracy %를 부여하지 않습니다. 품질은 언어, 오디오, speaker, vocabulary에 따라 달라집니다. 가격 예시는 2026년 8월 16일 확인했습니다.
짧은 결론: AI, 수동, hybrid?
| 상황 | 추천 시작점 | 이유 |
|---|---|---|
| 깨끗한 podcast/webinar/video | AI + 빠른 review | 빠른 draft와 집중 QA. |
| 노이즈 많은 multi-speaker interview | AI + 세심한 human review | 초기 작업은 절약되지만 어려운 구간은 검토 필요. |
| 법률·의료·high-stakes record | Human-reviewed workflow | 이름·숫자·진술 오류 비용이 큼. |
| 대규모 archive | AI-first | scale과 turnaround 우선, risk 기준 review. |
| 전문용어·고유명사 | AI + glossary/context + QA | context는 도움되지만 proper noun은 집중 확인. |
| 기밀 media | processing 방식에 따라 | 위치, retention, access, contract, human exposure를 확인. |
AI와 수동 전사는 단순 속도 차이가 아니라 workflow 차이
AI transcription은 ASR로 speech를 text로 변환합니다. 수동 전사는 사람이 듣고 해석하고 수정합니다. 많은 professional service는 ASR draft를 사람이 verify하는 hybrid 방식을 사용합니다.
품질은 marketing accuracy 숫자보다 녹음 조건에 더 크게 좌우됨
단일 “98%” 수치는 중요한 차이를 숨깁니다. language, accent, microphone, noise, overlap, names, technical vocabulary가 결과를 크게 바꿉니다. Whisper documentation도 language별 performance 차이를 명시하며 ASR output은 guaranteed ground truth가 아닙니다.
- clean close-mic speech는 AI와 사람 모두에게 유리.
- overlap은 text error와 speaker attribution error를 동시에 유발.
- names, numbers, products, technical terms는 별도 review.
- music, long silence, heavy editing은 segmentation에 영향.
- language/dialect coverage는 model별 차이.
- transcript가 맞아도 subtitle timing/speaker assignment는 별도 오류 가능.
모호한 accuracy보다 WER와 error type을 보세요
WER는 reference 대비 substitution, deletion, insertion을 계산합니다. CER는 character level error입니다. 실제 업무에서는 error 중요도가 더 중요할 수 있습니다. surname 하나나 decimal 하나의 오류가 여러 filler-word 차이보다 심각할 수 있습니다.
| Metric | 알 수 있는 것 | 알 수 없는 것 |
|---|---|---|
| WER | word-level error. | 각 error의 business impact. |
| CER | character-level error. | semantic/editorial importance. |
| Speaker accuracy | 올바른 speaker에 배정됐는지. | text 자체 정확성. |
| Subtitle QA | timing, segmentation, readability, overlap. | pure transcript quality. |
| Critical-field review | names, dates, numbers, sensitive terms. | 전체 fluency. |
AI와 사람이 어려워하는 조건
| 조건 | AI | 사람 |
|---|---|---|
| clean single speaker | 매우 좋은 AI use case. | 정확하지만 전부 수동이면 비효율적. |
| heavy noise | error가 크게 증가할 수 있음. | context가 도움되지만 들리지 않는 audio에는 한계. |
| overlapping speakers | text와 diarization 모두 실패 가능. | context는 유리하지만 simultaneous speech는 어렵다. |
| strong accent/dialect | model coverage에 따라 다름. | 익숙한 transcriber가 유리할 수 있음. |
| technical terminology | 그럴듯한 잘못된 term으로 대체 가능. | domain knowledge가 강점. |
| names/numbers | high-value error category. | source/context 확인 필요. |
| long repetitive archive | 매우 scalable. | manual-only는 느리고 비쌈. |
속도: first draft는 AI, final turnaround는 QA가 결정
AI는 보통 real-time listening보다 훨씬 빠르게 draft를 만듭니다. 그러나 generation time과 publish-ready time은 다릅니다. upload부터 final approval까지 측정해야 합니다.
비용: 분당 가격보다 전체 workflow를 비교
Rev는 현재 AI Transcription 0.25 USD/분, Human Transcription 1.99 USD/분부터 표시합니다. Happy Scribe는 AI minutes를 plan에 포함하고 Human Proofreading을 약 2 USD/분부터 제공하며 언어에 따라 더 높습니다. 이는 실제 예시이며 universal average가 아닙니다.
| 비용 | AI-first | Human-first |
|---|---|---|
| Initial transcription | 낮은 marginal cost, 높은 scale. | 높은 per-minute labor cost. |
| QA | clean audio에서는 적고 어려운 파일에서는 많음. | human process에 포함돼도 quality control 필요. |
| Speaker labels | 자동화 가능, review 필요. | manual이지만 context-aware. |
| Terminology | glossary/context와 correction 필요 가능. | specialist 필요 가능. |
| High volume | 경제적 장점이 큼. | human capacity가 bottleneck. |
| Critical errors | 저렴한 draft도 error 누락 시 비쌈. | 높은 초기 cost가 risk를 줄일 수 있음. |
Privacy는 AI vs 사람의 단순 이분법이 아니라 workflow 결정
confidentiality는 processing location, access, retention, contract에 따라 달라집니다. Cloud AI는 third-party infrastructure를 쓸 수 있고, human transcription은 사람이 content를 봅니다.
- processing location과 retention 확인.
- media/transcript가 training에 쓰이는지 확인.
- 필요 시 encryption 확인.
- human service에서 누가 access하는지 확인.
- regulated project는 구체적인 compliance scope 확인.
- Subvideo.ai는 Guest upload가 EU server에서 처리되고 48시간 후 자동 삭제된다고 안내합니다.
Human involvement를 늘려야 할 때
- transcript가 evidence/formal record인 경우.
- names, numbers, quotes, terms를 정확히 verify해야 하는 경우.
- difficult accents, crosstalk, poor audio.
- editorial nuance가 중요.
- domain expertise 필요.
- 숨은 error 비용이 review 비용보다 큰 경우.
AI-first가 적합한 경우
- 많은 시간의 media 처리.
- 빠른 searchable draft 필요.
- audio quality가 비교적 좋음.
- transcript가 subtitle/editing/search/repurposing의 중간 단계.
- high-risk section만 review 가능.
- multilingual processing을 scale해야 함.
2026 실용 workflow: AI first, risk가 높은 곳에 human effort
- 강한 ASR model로 transcript 생성.
- multi-speaker면 diarization을 별도 task로 처리.
- names/products/technical terms용 glossary/reference 사용.
- noise, overlap, fast speech, speaker change 우선 검토.
- names, numbers, dates, quotes 전용 pass 수행.
- subtitle에서는 text 이후 timing, segmentation, reading speed, line break 별도 검토.
- 정말 critical한 material만 deep human review.
- correction/minute와 total QA time 측정.
Subvideo.ai가 hybrid model에서 맡는 역할
Subvideo.ai는 outsourced human transcription보다 AI-first subtitle workflow에 맞춰져 있습니다. Whisper-based subtitles, optional speaker detection, visual timeline/editor manual correction, translation, SRT/VTT/ASS 또는 burned-in MP4 export를 제공합니다. Guest mode는 현재 EU에서 처리되고 48시간 후 삭제됩니다.
간단한 decision framework
Low risk + high volume
AI-first, sample과 critical field review.
Medium risk + difficult audio
AI draft + structured human QA.
High risk
Human-reviewed workflow와 review process 기록.
Subtitle 목적
text 이후 timing, speaker, readability 확인.
Sensitive media
location, retention, access, contract 기준 선택.
Cost 불확실
대표 batch에서 실제 QA minutes 측정.
AI는 human transcription을 없앤 것이 아니라 사람이 가장 큰 가치를 만드는 지점을 바꿨습니다. scalable media workflow에서는 AI-first가 효율적인 첫 단계가 되고, difficult/high-stakes content에서는 human review가 신뢰 가능한 final output을 만드는 quality layer입니다.
AI vs 수동 전사 FAQ
AI가 사람보다 정확한가요?
일률적으로 말할 수 없습니다. clean audio에서는 AI가 매우 강하고, context·difficult accent·terminology에서는 사람이 유리할 수 있습니다.
WER란?
reference 대비 substitution, deletion, insertion을 측정합니다. 낮을수록 좋지만 error의 중요도는 나타내지 않습니다.
수동 전사는 항상 99%인가요?
아닙니다. 일부 service가 조건 아래 보장하지만 audio, language, expertise, QA에 따라 달라집니다.
Human Review가 중요한 때는?
names, numbers, quotes, high-stakes facts가 정확해야 할 때입니다.
AI가 더 저렴한가요?
volume과 draft에서는 보통 그렇지만 editing, QA, storage, specialist review를 포함해 비교해야 합니다.
subtitle best workflow는?
AI 생성 → transcript 수정 → speaker 확인 → timing/segmentation/readability 검토입니다.